Why multimodal problems feel alive
I like multimodal problems because the evidence is rarely obedient. A silent but visually salient object can look convincing. A moving background can be distracting. A spoken or written reference can be underspecified.
That messiness is the point. When audio, vision, motion, and language disagree, the model has to reveal what it trusts. The interesting work is not only to fuse modalities, but to decide which evidence should matter for this particular case.
This is why audio-visual segmentation is such a good playground for me: the output is concrete, the failures are visible, and the research question stays close to the real scene.