cs.CV

It can look right and still have the wrong shape — a benchmark for surfaces beyond the observed time window

cs.CV Yukun Shi, Minglun Gong Jul 2026

Reconstruction of dynamic scenes is almost always evaluated inside the window that was filmed, yet real uses need the geometry at times after filming stopped. This work builds a controlled diagnostic benchmark with exact future ground truth and shows that looking right and having the right shape, beyond the window, do not move together.

Paper overview (our summary)

  • Field (arXiv category)cs.CV
  • AuthorsYukun Shi, Minglun Gong
  • Submitted2026-07-23
  • arXiv ID2607.21471v1

Key points

  • Rebuilding a moving scene is judged, as a rule, only within the stretch of time that was filmed, while real uses need the surface at moments the camera never saw.
  • The contribution is a deliberately narrow diagnostic suite: variety of scenes is given up in exchange for knowing the later shape exactly and for catching a measure that fires when it should not.
  • Methods train on the first 75 percent of a sequence and are scored on the held-out future by distance, with the future score itself as the primary number.
  • Motions that should show no gap are planted as falsification controls, and they behave as designed.
  • How good the later frames look and how correct the later shape is do not move together statistically, so the appearance-based numbers the field publishes say nothing about geometry beyond the window.

1Nobody was measuring outside the filmed window

There is a great deal of work on reconstructing moving scenes in three dimensions, and its evaluation almost always happens inside the window that was filmed: how faithfully the captured range was reproduced. What the actual uses need, though, is usually the shape after that.

Drawing something over real footage, having a robot touch an object, planning a motion in advance, all of these need the position of a surface at a time after filming stopped. No shared yardstick for that existed.

2What is measured, and how it can be falsified

  1. 1TrainLearn from the first 75 percent of the sequence, the observed part only
  2. 2Score the futureOn the held-out remainder, score the per-frame extracted surface by distance
  3. 3Diagnose the gapTake the future score itself as the primary number and read the gap against the observed part as a diagnostic
  4. 4Plant falsificationsMix in cases that should show no gap at all, such as motions that leave the surface unchanged

What matters in this design is the falsification controls. Mixing in conditions where no gap should appear reveals whether the measure captures a meaningful difference or simply assigns a difference to everything. In practice, the motion that leaves the surface unchanged is reported to show no gap. A study that builds a measure has also built the device for doubting it.

3Looks and shape turned out to be decoupled

Gap remaining even for futures predictable in principle2.7 to 4.1 timesFuture error on one backbone
Gap persisting on a second backbone2.0 to 6.6 timesBoth backbones share the same temporal mechanism
Records mentioning benchmarks among the 1,900 this site holds as of 2026-09-0459331.2% of the total

The heaviest finding is that how good the later frames look and how correct the later shape is do not move together in any statistical sense. The appearance-based numbers the field publishes say nothing about geometry beyond the window. Better-looking output is no guarantee that the shape is right. The error is also reported to be structured rather than scattered, concentrating where the surface moves.

4The eight moves, side by side

Across these articles, eight moves for turning an existing model to a new purpose have been examined. Laying out how often each appears in the records here shows where the center of gravity sits. The denominator is the 1,900 records this site holds as of 2026-09-04.

The moveWhere it intervenesRecords mentioning it, of those same 1,900
1 InsertLeave the model alone, strengthen what passes through21 plug-in (1.1%)
2 AddAdd adapters, preserve the past by distillation107 adapters and similar (5.6%)
3 Align premisesMake teacher and student see the same range89 distillation (4.7%)
4 Drop computationLook for redundancy in another direction127 reuse and redundancy (6.7%)
5 Move onto a baseDrop bespoke models for a foundation model89 foundation models (4.7%)
6 Make the dataIf ground truth is gone, synthesize the damage148 synthetic data (7.8%)
7 Remove the scaffoldingTake the devices down and build plainly152 pretraining from scratch (8.0%)
8 Measure againFind the failure that sits outside the yardstick593 benchmarks (31.2%)

The largest group is the eighth, the records concerned with measuring, at over three in ten. Records mentioning transfer, adaptation or fine-tuning number 359 (18.9%), more than twice the 152 (8.0%) about pretraining from scratch. The center of gravity has moved from building to reusing, and the largest share of effort now goes into measuring what the reuse produced.

Why it matters

However much a method is refined, an improvement cannot be confirmed if the quantity being measured is the wrong one. As long as appearance is what gets measured, errors in shape stay invisible. Deciding first what quantity the use actually needs, and planting conditions where no difference should appear, is a design that transfers to building any yardstick for evaluation.

FAQ

Why measure outside the filmed window?
Settings such as drawing graphics over live footage, having a robot make contact, or planning a movement before it happens all need where a surface will be after the camera stopped.
What is a falsification control?
A condition planted in advance where no difference should appear, which reveals whether a measure captures meaningful differences or assigns differences to everything.
Are appearance metrics not enough?
The paper reports that how good the later frames look and how correct the later shape is do not move together statistically, so looking better is no guarantee that the shape is right.

Sources (primary)

Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.

#AI research#arXiv#Computer vision#3D reconstruction#Benchmarks
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.