Multimodal Collaborative Agents for Next-Gen Commerce — Nidhi Kaushik Vyas, Google DeepMind
Asked to redo a living room on a budget, the agent Nidhi Kaushik Vyas walks through does not start recommending. It works out what it does not know, then picks the single question worth asking: how wide is the room. Everything downstream is moot if the furniture will not fit, so that one answer carries the most information per turn. Choosing the next question by expected information gain, rather than working down a checklist, is the heart of her argument about agents that handle fuzzy intent. Most shopping agents behave like a wrapper around the search bar, she says, assuming the user already holds a well formed goal and the vocabulary to express it. Real users arrive with a vibe, and closing that articulation gap is the agent's job rather than theirs. Her loop runs discovery, research, response. Discovery assembles a working state from past conversation, personal context and any reference images, separating the hard constraints in a query from the soft ones a picture only implies, attaching a confidence score to each, and flagging the variables that must refresh in real time because stale inventory makes an answer worthless. Research decides how to ask: for a subjective preference a visual board beats a text question, and hovers and clicks then feed the confidence model. Response picks a shape, a summary for a policy question, a comparison table for two products, imagery for style. Holding it together is an autorater at every stage, including one that flips part of a query to check the extracted constraints move when they should and stay put when they should not. Speaker info: - https://www.linkedin.com/in/nidhivyas/ Timestamps: 0:00 - Agents that behave like a wrapper around the search bar 1:08 - The articulation gap: users arrive with a vibe 2:03 - The loop: discovery, research, response 4:45 - Building a working state from images and context 6:33 - Variables that have to refresh in real time 7:28 - Grading the state: facts, calibration, counterfactuals 8:21 - The intent gap, and picking the highest value question 10:10 - Bridging a constraint to the product ontology 11:04 - When a visual board beats a text question 12:57 - Choosing the response format, and grading it 16:38 - Four takeaways 17:34 - Questions: merchant ontologies, UCP, and agent buyers




Join the discussion
Sign in to join the discussion
Sign in