Dataset profiles
nuPlan: Planning Data Requires More Than Logged Motion
A decision-focused review of closed-loop planning scenes and metrics: what the evidence can support, where teams can over-read it, and how to validate the next move.

Working hypothesis
nuPlan: Planning Data Requires More Than Logged Motion should be read as a testable operating claim, not a universal verdict. If planning data requires more than logged motion is the useful part of this collection, performance should improve or become more diagnosable on a target-domain bridge set built around the same conditions.
Evidence that would change our recommendation: Narrow the recommendation if the gain disappears on target sensors, if ontology mapping creates more error than the representation removes, or if the relevant operating slices are absent.
Start with the transfer question
A public dataset is a record of particular collection choices, not a promise about a future operating domain. For closed-loop planning scenes and metrics, the useful question is not whether the collection is respected or large. It is whether its sensors, scenes, labels and failure cases make it a credible bridge to the decision your model must make.
This field note examines planning data requires more than logged motion. The linked primary source remains authoritative for task definitions, collection details and reported results. Our focus is the product decision that follows: what must a team verify before transferring the data, benchmark result or ontology into its own program?
What to preserve from the collection
Begin with the acquisition and labeling contract. A benchmark score cannot tell you which deployment conditions were absent, whether clocks or coordinate frames were transformed, or how ambiguous examples were treated. Preserve the following evidence in the review record:
- Acquisition geography, hardware, viewpoint and environmental conditions.
- The label ontology, inclusion rules and known gaps in coverage.
- Sequence structure, sensor timing and transformations applied before release.
- A target-domain bridge set that exposes where benchmark and deployment distributions diverge.
These fields need not all become training targets. Some belong in provenance, slice metadata or the evaluation report. Their purpose is to explain why transfer succeeds in one condition and fails in another.
Where benchmark confidence can mislead
The central risk is that benchmark scale or familiarity hides geography, sensor and ontology mismatch. Familiarity can make a collection feel more representative than it is. Sensor placement, geography, environment, class definitions and sequence construction can each change the cost and meaning of an error.
Do not answer that concern with one larger aggregate. Break results down by the conditions that matter operationally, then compare them with a small target-domain bridge set. If the bridge set fails for ontology reasons, more benchmark pretraining will not repair the definition. If it fails for perception reasons, the error slices show which evidence to collect next.
A defensible evaluation plan
- Frame the decision. Write down the behavior the public data is meant to improve or test.
- Calibrate on ambiguity. Select ordinary scenes, hard cases and examples likely to differ from deployment.
- Diagnose disagreement. Review transfer errors by condition instead of one benchmark aggregate.
- Set the release gate. Keep a target-domain evaluation set separate from pretraining and the bridge set.
The output should be a short decision memo: which capabilities the collection can test, where it cannot support a claim, which target-domain evidence closes the gap, and what result would justify further investment.
Questions for the decision meeting
Before approving production work, the product, data and domain owners should be able to answer four questions in the same language:
- What changes if this works? Name the model behavior, evaluation decision or operational risk this evidence is meant to improve.
- Which conditions are still unrepresented? List the environments, sensors, actors, embodiments or failure modes that remain outside the claim.
- Where can qualified reviewers still disagree? Decide whether the remedy is more context, a clearer rule, an uncertainty label or domain adjudication.
- What result would stop or redirect the program? Define that threshold before scale, while the team can still change the collection and ontology inexpensively.
Write the answers into a one-page decision record and attach the calibration evidence. That record is more useful than a broad claim that the data is “high quality”: it identifies the intended use, the boundary of the evidence, the unresolved risks and the person accountable for accepting them. Revisit it when the model, ontology, collection hardware or operating environment changes.
What we would do next
We would select a compact, representative bridge set from the target environment and review every disagreement that could alter the operating decision. The resulting record becomes the first version of the ontology, calibration examples, escalation policy and acceptance test—not a polished demo disconnected from production.
The pilot should produce evidence even when the recommendation is to stop. A useful outcome may be a narrower ontology, a missing sensor requirement, a revised evaluation slice or proof that the proposed signal does not justify its cost. That is preferable to scaling a workflow whose assumptions have never been tested.
Sources and further reading
Primary references support the underlying dataset, method or release. The operational recommendations and limitations are Grasp's analysis.