cs.CV cs.LG

If the clean original is physically gone, synthesize the damage instead — a film-degradation pipeline and an 81,576-frame benchmark

cs.CV Mikołaj Jastrzębski, Dawid Glinkowski, Dawid Zieliński, et al. (6) Jul 2026

Restoring old film needs pairs of damaged and clean footage, yet the clean original is physically lost and cannot be recovered. This work models the analog-to-digital path as a composition of degradation families grounded in physics, and assembles an evaluation base of 81,576 frames taken from real archival footage.

Paper overview (our summary)

  • Field (arXiv category)cs.CV(+1)
  • AuthorsMikołaj Jastrzębski, Dawid Glinkowski, Dawid Zieliński, et al. (6)
  • Submitted2026-07-02
  • arXiv ID2607.02131v1

Key points

  • Restoring archival film needs damaged and clean pairs, but the clean original cannot be physically recovered.
  • The journey from film to digital file is treated as several distinct families of artifact combined under control: grain that follows the strength of the signal, scratches described by parameters, and camera movement that holds together frame to frame.
  • Splitting degradation into families makes the degree of damage controllable, so diverse regimes can be generated deliberately.
  • A curated evaluation base of 81,576 high-resolution frames from real archival footage is released alongside it.
  • Models trained with the pipeline generalized better to real footage, while the benchmark revealed systematic failure modes of current methods.

1The ground truth is physically gone

Teaching a machine to restore old film requires pairs: damaged footage alongside its clean version. Those pairs cannot be made in principle, because the original state of deteriorated film no longer physically exists and cannot be reshot. So the usual route has been to take clean footage and add artificial damage to it. The trouble was that the artificial damage did not resemble the real thing.

Film deterioration differs from simple noise in how the grain sits, how scratches run, and how the image wanders from frame to frame.

2Build the damage itself, grounded in physics

  1. 1Treat it as a processDecompose the path from analog to digital into a composition of degradation families
  2. 2Add the grainApply grain whose behavior depends on the strength of the signal
  3. 3Add the scratchesIntroduce scratches whose form is set by parameters
  4. 4Add the wanderGive camera motion that stays coherent across frames
  5. 5Vary the regimeRecombine these to generate controlled degrees of degradation

The point is that degradation is not treated as one lump of noise but split into families with different origins. Split apart, each component can be dialed up or down, so conditions close to a real archive can be aimed at deliberately. Including motion that stays coherent across frames matters for the same reason: damage applied independently per frame never resembles the real thing.

3Scale, and the yardstick for measuring

Frames in the curated archive set81,576Frames at high resolution, drawn from genuine archival material
Records mentioning synthetic data among the 1,900 this site holds as of 2026-09-041487.8% of the total
Records mentioning benchmarks, of those same 1,90059331.2% of the total

The other thing this work builds is a basis for evaluation. In a field without training data, evaluation data is usually missing too. Without gathering real footage into a shared yardstick, methods cannot be compared fairly. That just over three in ten of the 1,900 records this site holds as of 2026-09-04 mention benchmarks says something about how much of the field's effort now goes into measuring rather than building.

4The limits of learning from synthesis are shown too

Models trained with the pipeline are reported to generalize better to real-world footage, and the proposed benchmark is reported to reveal systematic failure modes of current methods. Learning from synthetic data means failures persist wherever the assumptions behind the synthesis do not hold. That the same authors also release the yardstick on which those failures become visible is the honest part of this work.

Why it matters

Where ground truth is physically gone, one route is to synthesize the damage rather than collect the truth. What matters then is splitting the assumptions into controllable families, and releasing alongside them a yardstick on which the places those assumptions fail become visible. The thinking carries beyond cultural preservation to any setting where the original of a record has been lost.

FAQ

Why can paired data not be made?
The clean original state of deteriorated film is physically lost and can neither be recovered nor reshot, which forces reliance on synthesis.
Why split degradation into families?
Grain, scratches and frame-to-frame wander have different origins. Handled separately, each can be dialed up or down so that conditions close to reality can be aimed at.
Why release a benchmark at the same time?
Fields without training data usually lack evaluation data too. Without a shared yardstick, methods cannot be compared fairly and claims of improvement cannot be checked.

Sources (primary)

Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.

#AI research#arXiv#Computer vision#Synthetic data#Archives
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.