WitneyWW's picture
Measured audio-edit findings: no non-speech removals found; correct the earlier speech claim; add scan manifest
27ef2af verified
|
Raw History Blame Contribute Delete
3.01 kB
metadata
title: InsAVE-80K Sample Viewer
emoji: 🎬
colorFrom: indigo
colorTo: pink
sdk: static
app_file: index.html
pinned: false
license: mit

InsAVE-80K β€” input / output sample viewer

Twenty-nine inspected pairs from InsAVE-80K, the instruction-based audio-video editing dataset released with InstructAV2AV (paper: arXiv:2605.18467):

  • 5 eval samples β€” eval/00000 … eval/00004

  • 16 add_and_remove samples β€” indices 00000, 00006, 00112, 00167, 00217, 00504, 00539, 00688, 00955, 01111 (person and object removals: a framed photograph, a screen overlay, a horse and rider) plus 10993, 11593, 14593, 15493, 16093, 20293 β€” cases where the edit removes the person's voice from the soundtrack as well. Every forward instruction in this category is a removal, so reading a card right-to-left gives the matching addition example.

    The audio edit is implicit: no instruction in the file mentions sound. A scan of 4,246 pairs (all 3,418 non-person removals with no quoted speech, plus 400-row controls) found:

    • Removing a speaking person does remove the speech β€” Whisper transcripts show the quoted <S>…<E> words present in the original (overlap 0.6–1.0) and gone from the target (0.0–0.17).
    • But it is not selective: the gain-cancelled speech-band change is +0.00 dB in both the speaking-person and silent-person pools. What differs is the level β€” the total-wipe rate goes from 3.7% to 10.7% (z = 3.85). The speech goes because the whole track drops out.
    • No non-speech audio removals were found. Non-person removals and a control pool of silent person removals are indistinguishable (14.0% vs 14.6% "selective" rate; βˆ’1.5 vs βˆ’1.6 dB median band loss), and 92% of all pairs have a regenerated soundtrack regardless of what was removed.

    The audio unchanged cards show this directly: a howling husky, a howling beagle, a helicopter and a broadcasting television, all removed from the picture with the audio provably intact.

  • 4 further train samples β€” 00000 from clone_id, clone_id_voice, clone_voice and general_editing

An "at a glance" table ranks every sample by how much of the frame the edit changed and how much speech-band energy it removed; the audio removed filter chip isolates the audio-visual removals. Each card shows the input (source clip) and the output (edited target clip) side by side, with the forward instruction and the reverse instruction, plus:

  • 6 evenly-spaced frames from each clip
  • an audio spectrogram for each clip (0–8 kHz)
  • an original-vs-target waveform overlay
  • a per-frame |original βˆ’ target| difference heat map
  • summary numbers: mean pixel Ξ”, fraction of pixels changed, p99 pixel Ξ”, audio envelope correlation

Clips are re-encoded to 352p H.264 / 96 kbps AAC for web playback; all measurements are computed from the original 1280Γ—704 files.