The solution
Off-road navigation lacks structure: there is no fixed vocabulary for what is traversable, and traversability depends on both the environment and the vehicle's own dynamics. Neither variable can be hand-labeled at scale, so traversability has to be learned from the embodiment's own experience.
Against the trend of multi-sensor rigs and computationally intensive inference platforms, FLINT is a 21.6M-parameter backbone, 38x smaller than a comparable foundation-model backbone, that uses a single RGB camera and runs at 14.7 FPS on CPU alone. It scored higher on held-out terrain probes and produced cheaper, more accurate costmaps than a deployed foundation-model system (WildOS) on 23 of 24 replayed field logs.
The authors compared self-supervised learning signals and deployed the resulting models in closed-loop field trials: the best self-supervised head reached 99% autonomy over the route, outperforming a human-label-trained baseline deployed live on the same course.
Why it worked
Self-supervision from the platform's own experience matches training data to the exact embodiment and terrain it will face.
A small backbone on one RGB camera cuts cost and latency without losing the decision-relevant signal.
Held-out terrain probes and replayed field logs tested generalization before live deployment.
Beating a human-label-trained baseline live shows the labels, not the architecture, were the bottleneck.
What can be applied
For embodied skills, the embodiment's own logged experience can supervise a model small enough for commodity hardware; heavy sensing and computing are not always necessary.
Aftermath
FLINT beat the deployed WildOS foundation-model system on 23 of 24 replayed field logs and reached 99% route autonomy in closed-loop trials, supporting the paper's conclusion that heavy sensing and computing are not necessary for traversability estimation.
FOLLOW THE EVIDENCE