All field notes

Autoresearch on robotics training data

As an AI-native team, we've been using coding agents to help fine-tune models for our own tasks. We wanted to share how HFlow fits into that work, and open source an example for anyone to explore.

HFlow is an open source SDK for data teams that collect, transform, and curate Physical AI data. In this workflow, we use HFlow to prepare and track data for an AI agent to train and evaluate models for use with HFlow's existing evaluation workflow.

Prepare the data before training

We use HFlow to process the source media and prepare it for training. It identifies duplicate images across the datasets, preserves where each sample came from, and creates reproducible training, validation, and test splits. This gives the agent a consistent dataset to work with while it explores changes to the training code.

In our own experiments, we had cases where the model's answers made us question whether we had converted the labels correctly. Before trying another training approach, we checked the original labels against the prepared examples, and checked that the images still matched their labels. The conversion checks passed. But the investigation also exposed contradictory source labels, duplicate data across releases, and gaps in how we recorded where the data came from.

These are different problems with different fixes. If we had kept asking the agent to improve training without checking the data, we would have been experimenting without understanding what the model was learning from. This is also why we preserve where every duplicate came from, rather than just keeping the first image and dropping the rest of its history.

In our open source example, we combine the datasets and deduplicate before creating the splits. Identical images with different encodings can still be identified by their decoded pixels. If duplicate images have conflicting labels, we exclude them and record the conflict instead of choosing a label for the agent.

Exact-image deduplication still has limits. Two neighbouring frames from the same recording may look almost identical without having identical pixels. If the source doesn't provide reliable recording identities, we can't claim that the splits contain independent recordings just because there are no exact duplicates.

Train and compare candidates

The usual model research loop is to train a baseline, change the training approach, and compare the result. If the data or evaluation changes between runs, we can't tell whether a better score came from the code change or from the experiment itself.

Andrej Karpathy's autoresearch project gives an agent a tight version of this loop. First, we measure the untouched baseline. Then the agent edits the training code, runs an experiment, and compares the candidate's score with the baseline. It keeps a useful change or tries another one. Unlike your traditional hyperparameter sweep, the agent can change the code itself.

To compare runs, we keep the data split, evaluation code, and base model fixed. The agent changes only the training file. Each candidate trains from the same base model, and a separate evaluator checks its saved adapter. We keep a record of every attempt, including runs that fail or run out of time.

The validation score guides the search. After the agent's attempts are over, we evaluate the selected candidate once on the held-out test partition. These limits on exact-image deduplication matter when we interpret both scores.

The public autoresearch workflow: HFlow prepares and deduplicates the data, an agent changes the training file within a fixed budget, and the operator freezes selection before one held-out confirmation.

What we learnt from our own experiments

Some of the useful lessons came from deciding what to investigate next, and whether a result was actually good enough to keep. An agent can make these iterations faster, but we still need to be clear about what we are asking it to improve.

A better score isn't always a better model

We had candidates that answered more examples correctly, but also made more unsupported guesses when the input was ambiguous. The overall score looked better, while a behaviour we cared about became worse.

If the workflow needs the model to avoid guessing, that needs to be part of the evaluation. We started looking at correct answers, unsupported answers, and how often the model answered separately. A model that refuses to answer everything isn't useful either, so simply reducing the number of wrong answers isn't enough.

The lesson for autoresearch is to decide what a useful improvement means before the agent starts. Otherwise, it can improve the metric we gave it while moving away from the behaviour we wanted. Keeping the individual predictions also helps us understand the tradeoff instead of just comparing two numbers.

A test set doesn't stay untouched by itself

We also had to think about evaluation across experiments. A run can keep its training and test rows separate, but if we look at the test failures and use them to decide what to train next, those results have already influenced the research.

We started separating the search from final evaluation more explicitly. The agent can inspect validation results and use them to propose the next change. Once we select a candidate, a separate evaluation checks it on the held-out data. If those results guide another round of research, we need to recognise that the data has become part of development.

This becomes more relevant as agents run more experiments. We can freeze the files and keep their hashes, but that doesn't make examples we've already learnt from untouched again.

Check what actually reaches the model

Another lesson was that using the same weights and images wasn't enough to make two evaluations equivalent. We observed differences when evaluating adapters before and after merging them into the base model, and when using different inference runtimes.

That led us to check the actual inputs, including image preprocessing and the rendered prompt, rather than just checking that the configuration looked the same. A change in the prompt format or image processing can affect the result even if we haven't changed the training.

For autoresearch, this means keeping the evaluation path consistent while the agent changes training. If we want to change the inference setup too, we treat that as a separate comparison. Otherwise, it becomes difficult to tell what caused an improvement.

Try the open source example

We kept the task small: count whether 0, 1, or 2 of the camera wearer's hands are visible in an image. It uses public Build AI evaluation releases and their published Gemini-generated labels. These are teacher labels, not human ground truth, and you don't need to call Gemini yourself to prepare them.

The example fine-tunes a tiny VLM with LoRA, which trains a small set of adapter parameters while keeping the base weights fixed. You can run the baseline and starter on CPU before giving an agent the training file to edit. It can explore sampling, adapter placement, loss weighting, or training-loop efficiency under the same time allowance. There isn't a particular coding agent built into the runner, and you can also make the changes manually.

HFlow handles manifest deduplication, provenance, and reproducible splitting. The example contains the training and evaluation runner. You can follow the whole process from preparing the public data to evaluating and exporting the selected model, without adopting our internal workflow.

CPU fine-tuning and short code trials: the earlier fixed-step pilot improved development macro-F1 from 0.1667 to 0.4889, while two short editable-code trials remained at 0.1667.

These are two separate experiments from the recorded CPU results. The earlier fine-tuning pilot shows room to improve the model, although it still failed on the one-hand class. The short editable-code trials check that the loop runs, and didn't improve the baseline. This isn't a single agent search curve or a held-out test result.

Coding agents are starting to take on more of the work in model research: proposing changes, running experiments, and learning from the results. Reliable data and evaluation make that work useful. HFlow helps teams prepare the data before the agent explores changes to the training code. Try our open source autoresearch example and let us know what you learn.