Hebbian Robotics builds fast, open-source SDKs for Physical AI data pipelines at scale. Robotics foundation model labs are buying massive volumes of egocentric video, and we help them filter and label the data that actually improves models.
Our flagship open source project is HFlow, data quality infrastructure for teams that collect, transform, and curate Physical AI data. Teams write quality checks as ordinary Python functions, while HFlow provides the durable execution, observability, and auditability around them.
We come from top embodied AI and ECE programs at Columbia University and the National University of Singapore. Our team has topped Stanford's LLM benchmarks and built infrastructure at Jane Street and Verkada, and we are backed by Y Combinator and angels from Oracle, Google DeepMind, and OpenAI.
Data quality is infrastructure
Quality control is still painstakingly manual, and metrics are rarely defined with enough rigor to be reproducible. We believe every result should be traceable across operators, sessions, and collection conditions.
Measure data before training
Today, we help teams evaluate data quality without training a robotics model. Data should be studied with the same seriousness and methodology that researchers apply to models.
Open source compounds
Data teams understand their collection systems and should own the checks that define quality. By building HFlow in public, an edge case caught by one team can become a reliable check for everyone.
Join us
If this thesis resonates with you, help us build the future of Physical AI data infrastructure. Explore HFlow or connect with Brandon and Kingston. Follow Hebbian Robotics at @hbr_pbc.