Abstract
Can a model distinguish between the sound of a spoon hitting a hardwood floorversus a carpeted one? Everyday object interactions produce sounds unique tothe objects involved. We introduce the sounding object detection task toevaluate a model's ability to link these sounds to the objects directlyinvolved. Inspired by human perception, our multimodal object-aware frameworklearns from in-the-wild egocentric videos. To encourage an object-centricapproach, we first develop an automatic pipeline to compute segmentation masksof the objects involved to guide the model's focus during training towards themost informative regions of the interaction. A slot attention visual encoder isused to further enforce an object prior. We demonstrate state of the artperformance on our new task along with existing multimodal action understandingtasks.