Hebbian Robotics

Pareto: high quality data curation for robotics

Brandon Ong
2026-07-16

Teams training AI models for robotics today are throwing away large portions of the data they collect when curating “high-quality” datasets. It is not only a painfully manual process. “High quality” is also hard to define, so most teams fall back on hand-crafted heuristics.

We built Pareto to help teams find their best training samples and run automated data consistency algorithms on them.

Our goal is to analyze and develop robotics datasets with the same rigor researchers bring to models. Upload a dataset, and Pareto indexes consistency metrics, identifies behavioral clusters, and surfaces potential anomalies. Now, instead of micromanaging the data collection process or applying brittle heuristics, we can apply algorithms that proxy what a consistent data collection process would produce.

One way we are algorithmically standardizing a dataset is through task velocity debiasing. This removes the artificial variation in your training data that comes from different human operators, or the same one on different days, executing the same motion at different speeds. On research benchmarks, this normalization technique has been shown to achieve up to a 60% improvement in task completion scores (Shi et al., 2025). Notably, this gain equals what scaling the training dataset by 2.5x achieves, underscoring the data efficiency of debiasing.

Pareto unlocks consistent data for more robust models, and more data-efficient training enables faster iteration cycles and lowers compute costs. We are excited to get the community's feedback on where we are headed. You can try Pareto in public beta.

We are gradually onboarding teams that find Pareto useful.