Download README.md from WitneyWW/insave-80k-sample-viewer: direct link, hf CLI and curl.
- Browser
- Download file 3.01 kB
-
https://huggingface.co/spaces/WitneyWW/insave-80k-sample-viewer/resolve/main/README.md
- Command line
-
hf download hf://spaces/WitneyWW/insave-80k-sample-viewer/README.md
-
curl -L -o README.md https://huggingface.co/spaces/WitneyWW/insave-80k-sample-viewer/resolve/main/README.md
title: InsAVE-80K Sample Viewer
emoji: π¬
colorFrom: indigo
colorTo: pink
sdk: static
app_file: index.html
pinned: false
license: mit
InsAVE-80K β input / output sample viewer
Twenty-nine inspected pairs from InsAVE-80K, the instruction-based audio-video editing dataset released with InstructAV2AV (paper: arXiv:2605.18467):
5 eval samples β
eval/00000β¦eval/0000416
add_and_removesamples β indices 00000, 00006, 00112, 00167, 00217, 00504, 00539, 00688, 00955, 01111 (person and object removals: a framed photograph, a screen overlay, a horse and rider) plus 10993, 11593, 14593, 15493, 16093, 20293 β cases where the edit removes the person's voice from the soundtrack as well. Every forward instruction in this category is a removal, so reading a card right-to-left gives the matching addition example.The audio edit is implicit: no instruction in the file mentions sound. A scan of 4,246 pairs (all 3,418 non-person removals with no quoted speech, plus 400-row controls) found:
- Removing a speaking person does remove the speech β Whisper transcripts show the quoted
<S>β¦<E>words present in the original (overlap 0.6β1.0) and gone from the target (0.0β0.17). - But it is not selective: the gain-cancelled speech-band change is +0.00 dB in both the speaking-person and silent-person pools. What differs is the level β the total-wipe rate goes from 3.7% to 10.7% (z = 3.85). The speech goes because the whole track drops out.
- No non-speech audio removals were found. Non-person removals and a control pool of silent person removals are indistinguishable (14.0% vs 14.6% "selective" rate; β1.5 vs β1.6 dB median band loss), and 92% of all pairs have a regenerated soundtrack regardless of what was removed.
The
audio unchangedcards show this directly: a howling husky, a howling beagle, a helicopter and a broadcasting television, all removed from the picture with the audio provably intact.- Removing a speaking person does remove the speech β Whisper transcripts show the quoted
4 further train samples β
00000fromclone_id,clone_id_voice,clone_voiceandgeneral_editing
An "at a glance" table ranks every sample by how much of the frame the edit changed and how much speech-band energy it removed; the audio removed filter chip isolates the audio-visual removals. Each card shows the input (source clip) and the output (edited target clip) side by side, with the forward instruction and the reverse instruction, plus:
- 6 evenly-spaced frames from each clip
- an audio spectrogram for each clip (0β8 kHz)
- an original-vs-target waveform overlay
- a per-frame
|original β target|difference heat map - summary numbers: mean pixel Ξ, fraction of pixels changed, p99 pixel Ξ, audio envelope correlation
Clips are re-encoded to 352p H.264 / 96 kbps AAC for web playback; all measurements are computed from the original 1280Γ704 files.