It can look right and still have the wrong shape — a benchmark for surfaces beyond the observed time window
Reconstruction of dynamic scenes is almost always evaluated inside the window that was filmed, yet real uses need the geometry at times after filming stopped. This work builds a controlled diagnostic benchmark with exact future ground truth and shows that looking right and having the right shape, beyond the window, do not move together.
Paper overview (our summary)
- Field (arXiv category)cs.CV
- AuthorsYukun Shi, Minglun Gong
- Submitted2026-07-23
- arXiv ID2607.21471v1
Key points
- Rebuilding a moving scene is judged, as a rule, only within the stretch of time that was filmed, while real uses need the surface at moments the camera never saw.
- The contribution is a deliberately narrow diagnostic suite: variety of scenes is given up in exchange for knowing the later shape exactly and for catching a measure that fires when it should not.
- Methods train on the first 75 percent of a sequence and are scored on the held-out future by distance, with the future score itself as the primary number.
- Motions that should show no gap are planted as falsification controls, and they behave as designed.
- How good the later frames look and how correct the later shape is do not move together statistically, so the appearance-based numbers the field publishes say nothing about geometry beyond the window.
1Nobody was measuring outside the filmed window
There is a great deal of work on reconstructing moving scenes in three dimensions, and its evaluation almost always happens inside the window that was filmed: how faithfully the captured range was reproduced. What the actual uses need, though, is usually the shape after that.
Drawing something over real footage, having a robot touch an object, planning a motion in advance, all of these need the position of a surface at a time after filming stopped. No shared yardstick for that existed.
2What is measured, and how it can be falsified
- 1TrainLearn from the first 75 percent of the sequence, the observed part only
- 2Score the futureOn the held-out remainder, score the per-frame extracted surface by distance
- 3Diagnose the gapTake the future score itself as the primary number and read the gap against the observed part as a diagnostic
- 4Plant falsificationsMix in cases that should show no gap at all, such as motions that leave the surface unchanged
What matters in this design is the falsification controls. Mixing in conditions where no gap should appear reveals whether the measure captures a meaningful difference or simply assigns a difference to everything. In practice, the motion that leaves the surface unchanged is reported to show no gap. A study that builds a measure has also built the device for doubting it.
3Looks and shape turned out to be decoupled
The heaviest finding is that how good the later frames look and how correct the later shape is do not move together in any statistical sense. The appearance-based numbers the field publishes say nothing about geometry beyond the window. Better-looking output is no guarantee that the shape is right. The error is also reported to be structured rather than scattered, concentrating where the surface moves.
4The eight moves, side by side
Across these articles, eight moves for turning an existing model to a new purpose have been examined. Laying out how often each appears in the records here shows where the center of gravity sits. The denominator is the 1,900 records this site holds as of 2026-09-04.
| The move | Where it intervenes | Records mentioning it, of those same 1,900 |
|---|---|---|
| 1 Insert | Leave the model alone, strengthen what passes through | 21 plug-in (1.1%) |
| 2 Add | Add adapters, preserve the past by distillation | 107 adapters and similar (5.6%) |
| 3 Align premises | Make teacher and student see the same range | 89 distillation (4.7%) |
| 4 Drop computation | Look for redundancy in another direction | 127 reuse and redundancy (6.7%) |
| 5 Move onto a base | Drop bespoke models for a foundation model | 89 foundation models (4.7%) |
| 6 Make the data | If ground truth is gone, synthesize the damage | 148 synthetic data (7.8%) |
| 7 Remove the scaffolding | Take the devices down and build plainly | 152 pretraining from scratch (8.0%) |
| 8 Measure again | Find the failure that sits outside the yardstick | 593 benchmarks (31.2%) |
The largest group is the eighth, the records concerned with measuring, at over three in ten. Records mentioning transfer, adaptation or fine-tuning number 359 (18.9%), more than twice the 152 (8.0%) about pretraining from scratch. The center of gravity has moved from building to reusing, and the largest share of effort now goes into measuring what the reuse produced.
Why it matters
However much a method is refined, an improvement cannot be confirmed if the quantity being measured is the wrong one. As long as appearance is what gets measured, errors in shape stay invisible. Deciding first what quantity the use actually needs, and planting conditions where no difference should appear, is a design that transfers to building any yardstick for evaluation.
FAQ
Why measure outside the filmed window?
What is a falsification control?
Are appearance metrics not enough?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2607.21471