Hi! I'm a Computer Science PhD student at Columbia University, advised by Vishal Misra and Dan Rubenstein.
I study how LLMs and agents use retrieved information: whether retrieval surfaces the right evidence,
and whether the model then uses that evidence correctly. The evidence that looks most relevant isn't
always the most useful, fair, or safe to surface. I build evaluations that pinpoint where these systems
fail, and use them to make retrieval-augmented models and agents more reliable.
Research interests: retrieval-augmented and agentic LLMs, agent memory, evaluation and interpretability,
and the risks retrieval creates, from bias to sensitive-information leakage.
We test whether dense retrievers respond to identity signals in queries (political ideology and dialect) across political news and health domains. Results: all five retrievers favor documents matching the query’s political lean and retrieve worse results for African American Language queries than for White Mainstream English, and these biases are encoded deep in the query embeddings rather than in surface vocabulary.
We propose ClusterSC to mitigate noise and the curse of dimensionality in disaggregate-level synthetic control by uncovering latent donor subgroups. Results: theoretical guarantees and significant MSE improvement on synthetic and real-world datasets.
We introduced an open-ended benchmark suite for embodied agent research, built on Minecraft and backed by a web-scale knowledge base. My role: built multimodal data pipeline for Minecraft Wiki and Reddit. I am highly grateful for this early project that inspired me to pursue agentic AI research.
Current Projects
MIDLLMAT: Mosaic Inference Defender Against LLM-Assisted Threats Andrew Tang, Matthew Connelly, Siddhartha Dalal, Vishal Misra IARPA BENGAL & NSF EAGER,
WIP, Updated September 2026
We measure the mosaic effect: whether sensitive facts can be rebuilt by aggregating individually less-sensitive documents, or recovered from an LLM’s own memory. We build retrieval pipelines that audit which facts in real declassified government records are already exposed, and that propose and verify redactions for documents under release review.
A lightweight dashboard that streams token logits, probabilities, and entropy during inference to spot when/where models learn concepts, experience mode collapse, or forget. We study links to curriculum design and catastrophic-drift debugging.