Mahsa Khoshnoodi
I work on multimodal AI, studying how models see and reason about the world, and building the tools that reveal when that understanding is real.
I study where and why vision-language models fail on fine-grained visual understanding. Rather than treating hallucination as an output artifact, I build diagnostic frameworks that trace it back to the exact point in a model's reasoning where perception breaks down. My longer view: capable multimodal systems will need a structured world model that connects seeing, understanding, and acting.
Perception to reasoning
I investigate how VLMs integrate visual and linguistic information to reach decisions. Even when models land on correct conclusions, their internal reasoning paths are often flawed or biased. I build interpretability tools that act as a microscope for AI, tracing information flow and exposing where perception fails to become genuine reasoning.
Diagnostic frameworks for VLMs
I develop evaluation frameworks that assess not just whether a model is correct, but whether its reasoning process is valid. Hallucination and bias appear heterogeneously across layers and architectures, so effective diagnosis means reading internal dynamics, not just observing outputs.
From seeing to acting
My long-term work targets systems that perceive, reason, and act reliably in the world. Drawing on structured world modeling and vision-language-action architectures, I aim to build multimodal systems that stay aligned with human values across the full loop from visual input to real-world decision.