Removing a sounding object from video and audio — 10 successes, 5 failures

Output of an agentic LangGraph pipeline run over source clips from JAVEdit-100k and InsAVE-80K. Only the source video and its soundtrack are borrowed: the pipeline captions the clip (Qwen3-Omni), picks what to remove (GPT-4o-mini), segments it (SAM3), erases it from the picture (EffectErase), verifies the erasure by re-segmenting, then separates its sound out of the mix (SAM-Audio, best of 10 candidates ranked with ImageBind). Play with sound on — the point of each pair is what you stop hearing.

60 clips run
23 passed (38.3%)
37 discarded
2×L40S per task · ~7–10 min/clip
7/28 audio selections repaired

Successes

Both modalities edited: the object is gone from the frame and its sound is gone from the mix, with the rest of the soundtrack still present. “Audio kept” is the residual RMS as a fraction of the source — high is good, it means only the object left, not the whole track. Four of these needed the corrected candidate selection described below.

passed InsAVE insave_06279
1280×704 · 23.98 fps · 3.42s · mp3 44kHz stereo
removed: mosquito chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “mosquito”
waveforms
source vs object-removed waveform
17.4%mask area (gate >15%)
+0.999visual removal (gate >0.40)
96.5%audio kept after removal
visualwinning prompt · seed 1132891577
9/10usable of 10 candidates
passedrescued by re-selection JAVEdit javedit_10d6a934a49d
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: woman chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “woman”
waveforms
source vs object-removed waveform
Original ranking picked a candidate keeping only 1.2% of the audio; the residual-energy guard picked one keeping 96.8%.
34.8%mask area (gate >15%)
+0.998visual removal (gate >0.40)
96.8%audio kept after removal
visualwinning prompt · seed 1132891577
4/10usable of 10 candidates
passedrescued by re-selection InsAVE insave_10090
1280×704 · 23.98 fps · 6.71s · mp3 44kHz stereo
removed: man chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “man”
waveforms
source vs object-removed waveform
Original ranking picked a candidate keeping only 25.2% of the audio; the residual-energy guard picked one keeping 94.2%.
26.1%mask area (gate >15%)
+1.000visual removal (gate >0.40)
94.2%audio kept after removal
textwinning prompt · seed 1778986134
2/10usable of 10 candidates
passedrescued by re-selection InsAVE insave_10590
1280×704 · 23.98 fps · 5.09s · mp3 44kHz stereo
removed: woman chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “woman”
waveforms
source vs object-removed waveform
Original ranking picked a candidate keeping only 21.0% of the audio; the residual-energy guard picked one keeping 89.5%.
34.4%mask area (gate >15%)
+1.000visual removal (gate >0.40)
89.5%audio kept after removal
textwinning prompt · seed 1453635084
6/10usable of 10 candidates
passedrescued by re-selection JAVEdit javedit_66311a9a9390
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: older man chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “older man”
waveforms
source vs object-removed waveform
Original ranking picked a candidate keeping only 0.4% of the audio; the residual-energy guard picked one keeping 67.3%.
39.9%mask area (gate >15%)
+1.000visual removal (gate >0.40)
67.3%audio kept after removal
textwinning prompt · seed 1335522078
7/10usable of 10 candidates
passed InsAVE insave_04135
1280×704 · 23.98 fps · 3.42s · mp3 44kHz stereo
removed: person chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “person”
waveforms
source vs object-removed waveform
26.0%mask area (gate >15%)
+1.000visual removal (gate >0.40)
103.9%audio kept after removal
visualwinning prompt · seed 1335522078
10/10usable of 10 candidates
passed JAVEdit javedit_0a0a140db4c0
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: man chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “man”
waveforms
source vs object-removed waveform
21.5%mask area (gate >15%)
+1.000visual removal (gate >0.40)
99.9%audio kept after removal
visualwinning prompt · seed 1335522078
10/10usable of 10 candidates
passed JAVEdit javedit_3eb55b3211e1
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: man chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “man”
waveforms
source vs object-removed waveform
35.8%mask area (gate >15%)
+1.000visual removal (gate >0.40)
94.5%audio kept after removal
visualwinning prompt · seed 1335522078
7/10usable of 10 candidates
passed InsAVE insave_04476
1280×704 · 23.98 fps · 6.71s · mp3 44kHz stereo
removed: man chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “man”
waveforms
source vs object-removed waveform
75.0%mask area (gate >15%)
+0.958visual removal (gate >0.40)
58.0%audio kept after removal
textwinning prompt · seed 1453635084
10/10usable of 10 candidates
passed JAVEdit javedit_1f3e390475b7
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: man chosen autonomously by Qwen3-Omni caption → GPT-4o-mini

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “man”
waveforms
source vs object-removed waveform
52.5%mask area (gate >15%)
+1.000visual removal (gate >0.40)
67.6%audio kept after removal
textwinning prompt · seed 1335522078
4/10usable of 10 candidates

Failures

One per rejection mode, plus the case that matters most: a clip the pipeline reports as a success whose audio is actually destroyed. Rejections are ordered cheapest-first in the graph, so most cost well under two minutes.

discarded JAVEdit javedit_047be501151f
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: woman chosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: inpainting did not remove the object

SAM3 was re-run on the inpainted clip and still found the target — the removal score is 1 − inpaint_area / original_area, so a negative value means the region got more detectable, not less.

visual_removal_check → visual_removal_score_below_threshold

SOURCE

source frames

INPAINTED · video only

result frames
49.8%mask area (gate >15%)
-0.382visual removal (gate >0.40)
discarded JAVEdit javedit_a0c929a7a33d
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: piano chosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: SAM3 could not find the object in frame 1

A fail-fast before segmenting the whole clip: the named target's first-frame mask covered less than 5% of the frame. Cheapest rejection in the graph (~40 s) and the most common one.

mask_check → first_frame_mask_below_threshold

SOURCE

source frames

NO ARTIFACT

Rejected before any mask or video was produced.
0.0%mask area (gate >15%)
discarded JAVEdit javedit_0f0dac495ba8
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: chosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: no sounding object in the caption

Qwen3-Omni captioned the clip and GPT-4o-mini found nothing in it that plausibly makes a sound, so there was nothing to remove. The clip never reaches a GPU-heavy stage.

sounding_object_extraction → no_sounding_object_found

SOURCE

source frames

NO ARTIFACT

Rejected before any mask or video was produced.
discarded JAVEdit javedit_a0d69d8e1c8d
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: person chosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: mask too small over the whole clip

The object was found but averages under 15% of the frame — too small for the inpainter to be worth ~10 minutes.

mask_check → mask_area_below_threshold

SOURCE

source frames

SAM3 MASK (last artifact)

The pipeline stopped here.
14.6%mask area (gate >15%)
passed the gates, unusable audio JAVEdit javedit_2788366a2890
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: young woman chosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Reached paired output, but the audio collapsed

Every gate passed — including audio_removal_check, which is still a mock that always returns 0.90. The object-removed audio retains just 8.9% of the source, i.e. near-silence. None of the ten SAM-Audio candidates was selective enough to rescue it, so this is a genuine failure the pipeline currently reports as a success.

SOURCE

source frames

RESULT · object removed from video + audio

result frames
source spectrogram
source audio
residual spectrogram
after removing “young woman”
waveforms
source vs object-removed waveform
32.3%mask area (gate >15%)
+1.000visual removal (gate >0.40)
8.9%audio kept after removal
textwinning prompt · seed 1335522078
0/10usable of 10 candidates

Two findings from the run

where clips diedncost
SAM3 first-frame mask too small21~40 s – 2 min
inpainting failed the re-segmentation check7~6–8 min
no sounding object in the caption6seconds
mask too small across the clip3~2 min

The visual gate is bimodal, not marginal. Median removal score among clips that reached it is 1.000, with 23 of 30 above 0.80 and a minimum of −5.9. EffectErase either erases the object completely or leaves it there — so the permissive 0.40 threshold used for this run bought nothing, and 0.80 would have cost no yield.

The best-of-10 audio ranking selected for silence. Candidates were ranked by maximum ImageBind text↔target similarity — but a separation that dumps the entire soundtrack into the target scores highest on exactly that metric, leaving a silent residual. One in four selections was degenerate this way. Filtering out candidates whose residual retains under 30% of the source, then ranking the survivors unchanged, repaired 7 of them offline with no GPU time (all ten candidate pairs are kept on disk). Three clips had no usable candidate at all — one is shown in the failures above. The pipeline's own audio_removal_check is still a stub that returns 0.90, which is why every one of these was reported as a pass.