Off-road navigation lacks structure: there is no fixed vocabulary for what is traversable, and traversability depends on both the environment and the vehicle's own dynamics. Neither variable can be hand-labeled at scale, so traversability has to be learned from the embodiment's own experience.

Against the trend of multi-sensor rigs and computationally intensive inference platforms, FLINT is a 21.6M-parameter backbone, 38x smaller than a comparable foundation-model backbone, that uses a single RGB camera and runs at 14.7 FPS on CPU alone. It scored higher on held-out terrain probes and produced cheaper, more accurate costmaps than a deployed foundation-model system (WildOS) on 23 of 24 replayed field logs.

The authors compared self-supervised learning signals and deployed the resulting models in closed-loop field trials: the best self-supervised head reached 99% autonomy over the route, outperforming a human-label-trained baseline deployed live on the same course.

Self-supervision from the platform's own experience matches training data to the exact embodiment and terrain it will face.

A small backbone on one RGB camera cuts cost and latency without losing the decision-relevant signal.

Held-out terrain probes and replayed field logs tested generalization before live deployment.

Beating a human-label-trained baseline live shows the labels, not the architecture, were the bottleneck.

For embodied skills, the embodiment's own logged experience can supervise a model small enough for commodity hardware; heavy sensing and computing are not always necessary.

FLINT beat the deployed WildOS foundation-model system on 23 of 24 replayed field logs and reached 99% route autonomy in closed-loop trials, supporting the paper's conclusion that heavy sensing and computing are not necessary for traversability estimation.

FOLLOW THE EVIDENCE

The sources

  1. FLINT: Fast Lightweight Inference for Traversability arxiv.org