Siltframe Get a fix →

Blog / 2026-10-03

Why off-road models call tree canopy “sky” — and why more augmentation didn’t fix it

We added synthetic bad weather to the training data and watched real-weather sky accuracy get worse on every architecture we tried. Chasing that down turned out to be the most useful thing we have measured, and it changed what we think the product is.

1 · The symptom nobody wants

We had been generating physically-modelled dust, night, fog, lens rain and lens mud, and fine-tuning segmentation models on them. On synthetic held-out tests it worked. Then we scored the same checkpoints on IDD-AW — real rain, fog, snow and low light, a dataset we never train on — and found that our augmentation improved drivable-surface detection while dropping sky accuracy on all five architectures, by 1 to 13 points. Vegetation fell too on the smaller models.

The convenient response would have been to report the drivable-surface win and keep quiet. Instead we went looking for the frames where it happened.

2 · What the failures actually looked like

The errors were not spread evenly. They concentrated on forest roads: tracks running under a closed tree canopy, where bright sky shows through leaves in patches. The models were labelling the whole upper half of those frames “sky”, canopy included.

Before: large parts of the tree canopy marked as sky
Trained on open terrain · magenta = “sky” · 80% of the tree pixels
After: the canopy is recognised as trees
With forest-road frames in training · 0%

One real frame from a drive neither model was trained on. Both models had the same weather augmentation; only the training data differs. Brightened for display — the models saw the original.

3 · The cause was in the training set all along

RELLIS-3D, the dataset most of this work is built on, was recorded on open terrain: fields, open trails, wide horizon. It contains very little of the one geometry that matters here — a road with trees closing over the top of it. A model trained on it learns a shortcut that is true in its world and false in a forest: bright, above the horizon, low texture → sky.

Degrading those same open-terrain frames with fog and darkness does not teach the model what a canopy is. It makes the shortcut more attractive: haze and night both wash out exactly the texture that would have given the leaves away. That is why augmentation made this particular failure worse rather than better.

4 · Fixing the generator recovered about half

Two of those losses were ours to fix, and we fixed them. Our fog and dust models were brightening sky pixels using a depth map that treats sky as a finite distance, and they were painting a veil over pixels that, physically, the camera could no longer see at all. So the generator became label-aware: sky is read from the label rather than guessed from depth, and pixels behind a veil dense enough to block more than 97 % of the light are marked void and simply not scored. That recovered roughly half of the sky and vegetation loss.

Half. After fixing the thing we were doing wrong. That was the clue that the rest was not an augmentation problem at all.

5 · The other half was data the model had never seen

We added 1,500 real frames from GOOSE — a German off-road dataset full of forest tracks, gravel and paved roads — and retrained. Same augmentation on both sides, same schedule, same seeds. Only the coverage changed.

Real adverse weather (IDD-AW)sky IoUvegetation IoUmean
RELLIS only+ GOOSERELLIS only+ GOOSERELLIS only+ GOOSE
Mask2Former (Swin-T)82.689.452.681.449.177.9
SegFormer-B072.087.943.978.141.070.0
SegFormer-B278.290.444.481.440.475.1
DeepLabV3-MobileNetV370.487.841.775.834.864.4
Identical fine-tunes with Siltframe augmentation, single seed, scored on real rain / fog / snow / low light. GOOSE: Fraunhofer IOSB, CC BY-SA 4.0.

Sky came back by 7–17 points, vegetation by 29–37, and the overall real-weather mean by 29–35 points. On the single forest-road frame above, the share of tree pixels called sky went from 80% to 0%.

For comparison: the best augmentation result we have ever measured on real adverse weather is worth a few points. The missing data was worth an order of magnitude more.

6 · The uncomfortable part

Once GOOSE was in the training set, we re-ran every augmentation arm on top of it — ours, off-the-shelf albumentations, and a combination — across four architectures. On real weather they all landed within ±1 point of the control. Our own augmentation still adds 3.5–5.2 points on synthetic held-out degradations, but on real frames, with the right data already present, it is within noise.

We sell synthetic data. This result says synthetic data is not the main lever, and we would rather publish that than find out from a customer. It is also why the offer is built the way it is: find the missing condition first, cover it with real data where real data exists, and use synthesis for the conditions you genuinely cannot capture — then judge the whole thing on real frames.

7 · How to find your own version of this

The canopy shortcut is specific to RELLIS-3D. Yours will be something else: a surface, a light condition, a machine, a season that your capture campaign happened to miss. It will not show up in your validation set, because your validation set came from the same campaign.

The cheapest way to start looking is to run a model you already have against conditions it has never seen and read the per-class failures rather than the headline number. We publish a free stress test that does exactly that — 165 labelled frames under dust, night, fog, lens rain and lens mud, with the scoring script: github.com/egeizgi/siltframe-stress-test. About ten minutes, no sign-up.

Limitations

  • Single seed per arm on the GOOSE runs, so treat 1–2 point differences as noise. The 26–35 point coverage effect is far outside that.
  • IDD-AW is Indian road scenes, not off-road terrain. It is the most honest real adverse-weather test we have access to, but it is a cross-domain test and the absolute numbers are low. What it measures reliably is the difference between training arms.
  • The canopy frame is one frame, chosen as the clearest example of a failure mode we found across many. It illustrates the effect; the table measures it.
  • GOOSE is one dataset from one country. “Add the missing coverage” is easy to write and expensive to do when the missing condition is not already in a public dataset — which is exactly the case for most machines.

which condition is your data missing?

We run this analysis on your model and your frames, and tell you what to capture or generate next.

Send a brief →