Removing a sounding object from video and audio — 10 successes, 5 failures
Output of an agentic LangGraph pipeline run over source clips from
JAVEdit-100k and InsAVE-80K. Only the source video and its soundtrack are borrowed:
the pipeline captions the clip (Qwen3-Omni), picks what to remove (GPT-4o-mini), segments it (SAM3),
erases it from the picture (EffectErase), verifies the erasure by re-segmenting, then separates its
sound out of the mix (SAM-Audio, best of 10 candidates ranked with ImageBind).
Play with sound on — the point of each pair is what you stop hearing.
60clips run
23passed (38.3%)
37discarded
2×L40Sper task · ~7–10 min/clip
7/28audio selections repaired
Successes
Both modalities edited: the object is gone from the frame and its sound is gone
from the mix, with the rest of the soundtrack still present. “Audio kept” is the residual
RMS as a fraction of the source — high is good, it means only the object left, not the whole
track. Four of these needed the corrected candidate selection described below.
passedInsAVEinsave_06279
1280×704 · 23.98 fps · 3.42s · mp3 44kHz stereo
removed: mosquitochosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “mosquito”
source vs object-removed waveform
17.4%mask area (gate >15%)
+0.999visual removal (gate >0.40)
96.5%audio kept after removal
visualwinning prompt · seed 1132891577
9/10usable of 10 candidates
passedrescued by re-selectionJAVEditjavedit_10d6a934a49d
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: womanchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “woman”
source vs object-removed waveform
Original ranking picked a candidate keeping only 1.2% of the audio; the residual-energy guard picked one keeping 96.8%.
34.8%mask area (gate >15%)
+0.998visual removal (gate >0.40)
96.8%audio kept after removal
visualwinning prompt · seed 1132891577
4/10usable of 10 candidates
passedrescued by re-selectionInsAVEinsave_10090
1280×704 · 23.98 fps · 6.71s · mp3 44kHz stereo
removed: manchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “man”
source vs object-removed waveform
Original ranking picked a candidate keeping only 25.2% of the audio; the residual-energy guard picked one keeping 94.2%.
26.1%mask area (gate >15%)
+1.000visual removal (gate >0.40)
94.2%audio kept after removal
textwinning prompt · seed 1778986134
2/10usable of 10 candidates
passedrescued by re-selectionInsAVEinsave_10590
1280×704 · 23.98 fps · 5.09s · mp3 44kHz stereo
removed: womanchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “woman”
source vs object-removed waveform
Original ranking picked a candidate keeping only 21.0% of the audio; the residual-energy guard picked one keeping 89.5%.
34.4%mask area (gate >15%)
+1.000visual removal (gate >0.40)
89.5%audio kept after removal
textwinning prompt · seed 1453635084
6/10usable of 10 candidates
passedrescued by re-selectionJAVEditjavedit_66311a9a9390
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: older manchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “older man”
source vs object-removed waveform
Original ranking picked a candidate keeping only 0.4% of the audio; the residual-energy guard picked one keeping 67.3%.
39.9%mask area (gate >15%)
+1.000visual removal (gate >0.40)
67.3%audio kept after removal
textwinning prompt · seed 1335522078
7/10usable of 10 candidates
passedInsAVEinsave_04135
1280×704 · 23.98 fps · 3.42s · mp3 44kHz stereo
removed: personchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “person”
source vs object-removed waveform
26.0%mask area (gate >15%)
+1.000visual removal (gate >0.40)
103.9%audio kept after removal
visualwinning prompt · seed 1335522078
10/10usable of 10 candidates
passedJAVEditjavedit_0a0a140db4c0
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: manchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “man”
source vs object-removed waveform
21.5%mask area (gate >15%)
+1.000visual removal (gate >0.40)
99.9%audio kept after removal
visualwinning prompt · seed 1335522078
10/10usable of 10 candidates
passedJAVEditjavedit_3eb55b3211e1
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: manchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “man”
source vs object-removed waveform
35.8%mask area (gate >15%)
+1.000visual removal (gate >0.40)
94.5%audio kept after removal
visualwinning prompt · seed 1335522078
7/10usable of 10 candidates
passedInsAVEinsave_04476
1280×704 · 23.98 fps · 6.71s · mp3 44kHz stereo
removed: manchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “man”
source vs object-removed waveform
75.0%mask area (gate >15%)
+0.958visual removal (gate >0.40)
58.0%audio kept after removal
textwinning prompt · seed 1453635084
10/10usable of 10 candidates
passedJAVEditjavedit_1f3e390475b7
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: manchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
SOURCE
RESULT · object removed from video + audio
source audioafter removing “man”
source vs object-removed waveform
52.5%mask area (gate >15%)
+1.000visual removal (gate >0.40)
67.6%audio kept after removal
textwinning prompt · seed 1335522078
4/10usable of 10 candidates
Failures
One per rejection mode, plus the case that matters most: a clip the pipeline
reports as a success whose audio is actually destroyed. Rejections are ordered cheapest-first in the
graph, so most cost well under two minutes.
discardedJAVEditjavedit_047be501151f
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: womanchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: inpainting did not remove the object
SAM3 was re-run on the inpainted clip and still found the target — the removal score is 1 − inpaint_area / original_area, so a negative value means the region got more detectable, not less.
removed: pianochosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: SAM3 could not find the object in frame 1
A fail-fast before segmenting the whole clip: the named target's first-frame mask covered less than 5% of the frame. Cheapest rejection in the graph (~40 s) and the most common one.
mask_check → first_frame_mask_below_threshold
SOURCE
NO ARTIFACT
Rejected before any mask or video was produced.
0.0%mask area (gate >15%)
discardedJAVEditjavedit_0f0dac495ba8
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: —chosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: no sounding object in the caption
Qwen3-Omni captioned the clip and GPT-4o-mini found nothing in it that plausibly makes a sound, so there was nothing to remove. The clip never reaches a GPU-heavy stage.
removed: personchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Rejected: mask too small over the whole clip
The object was found but averages under 15% of the frame — too small for the inpainter to be worth ~10 minutes.
mask_check → mask_area_below_threshold
SOURCE
SAM3 MASK (last artifact)
The pipeline stopped here.
14.6%mask area (gate >15%)
passed the gates, unusable audioJAVEditjavedit_2788366a2890
1280×720 · 25 fps · 4.86s · aac 16kHz mono
removed: young womanchosen autonomously by Qwen3-Omni caption → GPT-4o-mini
Reached paired output, but the audio collapsed
Every gate passed — including audio_removal_check, which is still a mock that always returns 0.90. The object-removed audio retains just 8.9% of the source, i.e. near-silence. None of the ten SAM-Audio candidates was selective enough to rescue it, so this is a genuine failure the pipeline currently reports as a success.
SOURCE
RESULT · object removed from video + audio
source audioafter removing “young woman”
source vs object-removed waveform
32.3%mask area (gate >15%)
+1.000visual removal (gate >0.40)
8.9%audio kept after removal
textwinning prompt · seed 1335522078
0/10usable of 10 candidates
Two findings from the run
where clips died
n
cost
SAM3 first-frame mask too small
21
~40 s – 2 min
inpainting failed the re-segmentation check
7
~6–8 min
no sounding object in the caption
6
seconds
mask too small across the clip
3
~2 min
The visual gate is bimodal, not marginal. Median removal score among clips
that reached it is 1.000, with 23 of 30 above 0.80 and a minimum of −5.9. EffectErase
either erases the object completely or leaves it there — so the permissive 0.40 threshold used
for this run bought nothing, and 0.80 would have cost no yield.
The best-of-10 audio ranking selected for silence. Candidates were ranked by
maximum ImageBind text↔target similarity — but a separation that dumps the entire
soundtrack into the target scores highest on exactly that metric, leaving a silent residual. One in
four selections was degenerate this way. Filtering out candidates whose residual retains under 30%
of the source, then ranking the survivors unchanged, repaired 7 of them offline with no GPU time
(all ten candidate pairs are kept on disk). Three clips had no usable candidate at all — one is
shown in the failures above. The pipeline's own audio_removal_check is still a stub that
returns 0.90, which is why every one of these was reported as a pass.