Andrew Tang

Hi! I'm a Computer Science PhD student at Columbia University, advised by Vishal Misra and Dan Rubenstein. I study how LLMs and agents use retrieved information: whether retrieval surfaces the right evidence, and whether the model then uses that evidence correctly. The evidence that looks most relevant isn't always the most useful, fair, or safe to surface. I build evaluations that pinpoint where these systems fail, and use them to make retrieval-augmented models and agents more reliable.

Research interests: retrieval-augmented and agentic LLMs, agent memory, evaluation and interpretability, and the risks retrieval creates, from bias to sensitive-information leakage.

Email  /  Google Scholar  /  X  /  LinkedIn

profile photo

Publications

project image Retrieval Sensitivity to Identity Signals in Queries
Andrew Tang, Nicholas Deas, Kathleen McKeown, Vishal Misra
EMNLP, 2026
[arxiv]

We test whether dense retrievers respond to identity signals in queries (political ideology and dialect) across political news and health domains. Results: all five retrievers favor documents matching the query’s political lean and retrieve worse results for African American Language queries than for White Mainstream English, and these biases are encoded deep in the query embeddings rather than in surface vocabulary.

project image ClusterSC: Advancing Synthetic Control with Donor Selection
Saeyoung Rho, Andrew Tang, Noah Bergam, Rachel Cummings, Vishal Misra
AISTATS, 2025
[arxiv] [code]

We propose ClusterSC to mitigate noise and the curse of dimensionality in disaggregate-level synthetic control by uncovering latent donor subgroups. Results: theoretical guarantees and significant MSE improvement on synthetic and real-world datasets.

MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
Linxi "Jim" Fan, Guanzhi Wang*, Yunfan Jiang*, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu†, Anima Anandkumar†
NeurIPS Datasets and Benchmarks Track, 2022  (Outstanding Paper Award)
[arxiv] [website] [code]

We introduced an open-ended benchmark suite for embodied agent research, built on Minecraft and backed by a web-scale knowledge base.
My role: built multimodal data pipeline for Minecraft Wiki and Reddit. I am highly grateful for this early project that inspired me to pursue agentic AI research.




Current Projects

project image MIDLLMAT: Mosaic Inference Defender Against LLM-Assisted Threats
Andrew Tang, Matthew Connelly, Siddhartha Dalal, Vishal Misra
IARPA BENGAL & NSF EAGER, WIP, Updated September 2026

We measure the mosaic effect: whether sensitive facts can be rebuilt by aggregating individually less-sensitive documents, or recovered from an LLM’s own memory. We build retrieval pipelines that audit which facts in real declassified government records are already exposed, and that propose and verify redactions for documents under release review.

project image TokenProbe: Visualizing LLM Learning in Real-Time
Andrew Tang, Amy Wu, Rashfiqur Rahman, Charlie Kerfoot, Vishal Misra
WIP, Updated June 2025
[website]

A lightweight dashboard that streams token logits, probabilities, and entropy during inference to spot when/where models learn concepts, experience mode collapse, or forget. We study links to curriculum design and catastrophic-drift debugging.





Last updated September 2026. Design and source code from Jon Barron's website