Dat Nguyen

Post‑doctoral Fellow in Computer Science, Harvard SEAS  ·  Basis Research Institute

profile.jpg

Short bio

I am a Joint Postdoctoral Fellow at Harvard’s Programming Languages and Formal Methods groups and the Basis Research Institute.

I am boardly interested in the the modeling of how we perceive the world, and the modeling of reasoning processes. To support this goal, I work in the emerging area between programming language, machine learning, and probabilistic programming language.

At Basis, I work on modeling uncertainty in symbolic world models in MARA, designing a robot design language that captures both morphology and control in R-ADA, and modeling LLM generation as an effect, as a framework for building agent harnesses, in effectful. At Harvard, I work on proof automation in Lean and causal systems for drug repurposing.

I completed my PhD doing machine learning and program synthesis-based debugging, and previously worked on computer vision problems: visually rich document information extraction.

Research interests

  • World models: learning, evaluation, and uncertainty
  • Languages and abstractions for robot design and LLM agents
  • Program synthesis and probabilistic programming
  • Neuro-symbolic systems modeled with LLMs, PPLs, and NNs
  • Reliable, explainable ML for software, including graph-based learning for code and documents

News

  1. Delivered a talk on WorldTest and AutumnBench at TAIC’26 (Thinking about AI’s Capability), a pre-ICML workshop at GIST. Check out the slides here and the blogpost here.
  2. My paper, ESC: Emotional Self-Correction for Reliable Vision-Language Models, is accepted at ECCV’26. arXiv, project.
  3. Awarded Gold Reviewer at ICML’26. Thanks to the area chairs and to the authors whose submissions were a pleasure to read.
  4. Our work, WorldTest, is accepted at ICML! WorldTest formulates world-model learning evaluation with environment-level queries that pose general questions about the environments, and we instantiated it with AutumnBench. See you in Korea! arXiv, project.
  5. Preprint, follow-up to NeuroSymbolicDG. We re-formulated image classification as spatial predicate induction over learned image primitives! arXiv.

Technical blogs

Project demos

NeuroSymbolicDG NeuroSymbolicDG
Domain-invariant classifier head for fine-grained bird recognition, via a PCFG over spatial layouts.
code · paper · blog · checkpoints
AutumnBench environments running live Autumn.cpp (ICML '26)
Autumn interpreter in C++. Powers MARA and AutumnBench. Try it live in the playground.
code · AutumnBench paper · blog · playground
ExoPredicator ExoPredicator (ICLR '26)
Learning abstract models of dynamic worlds for robot planning.
paper · openreview
VRDSynth VRDSynth (ISSTA '24)
Program synthesis for multilingual document information extraction.
code · paper
VirDA VirDA (TMLR '25)
Unsupervised domain adaptation by reusing the backbone with visual reprogramming.
code · paper
GNNInfer GNNInfer (ICSE '22, arXiv '24)
Inferring properties of graph neural networks.
paper
FFL FFL (ICSME '22)
Fine-grained fault localization for student programs.
code · paper

Positions

2025 to present
Joint Post-doctoral Fellow
2021 to 2024
PhD, School of Computing & Information Systems
University of Melbourne · Melbourne Research Scholarship
2016 to 2021
AI Research Engineer
Cinnamon AI Lab

Selected Publications

  1. ICML '26Benchmarking World-Model Learning with Environment-Level Queries. Archana Warrier, Dat Nguyen, Michelangelo Naim, Moksh Jain, Yichao Liang, Karen Schroeder, Cambridge Yang, Joshua B. Tenenbaum, Sebastian Vollmer, Kevin Ellis, Zenna Tavares.
  2. ICLR '26ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning. Yichao Liang, Dat Nguyen, Cambridge Yang, Tianyang Li, Joshua B. Tenenbaum, Carl Edward Rasmussen, Adrian Weller, Zenna Tavares, Tom Silver, Kevin Ellis.
  3. arXiv '25A Systematic Survey on Debugging Techniques for Machine Learning Systems. Dat Nguyen, Haoye Tian, Bach Le, Patanamon Thongtanunam, Shane McIntosh.
  4. arXiv '24Inferring Properties of Graph Neural Networks. Dat Nguyen, Hieu M. Vu, Cong-Thanh Le, Bach Le, David Lo, ThanhVu Nguyen, Corina Pasareanu.
  5. ISSTA '24VRDSynth: Synthesizing Programs for Multilingual Visually Rich Document Information Extraction. Dat Nguyen, Tung Do-Viet, Hung Nguyen-Duy, Tuan-Hai Luu, Hung Le, Bach Le, Patanamon Thongtanunam.
  6. arXiv '24Combining Induction and Transduction for Abstract Reasoning. Wen-Ding Li, Keya Hu, Carter Larsen, Yuqing Wu, Simon Alford, Caleb Woo, Spencer M. Dunn, Hao Tang, Michelangelo Naim, Dat Nguyen, Wei-Long Zheng, Zenna Tavares, Yewen Pu, Kevin Ellis.
  7. arXiv '23Adversarial Attacks on Code Models with Discriminative Graph Patterns. Dat Nguyen, Yang Zhou, Xuan Bach D. Le, Patanamon Thongtanunam, David Lo.
  8. ICSME '22FFL: Fine grained Fault Localization for Student Programs via Syntactic and Semantic Reasoning. Dat Nguyen, Thanh Le-Cong, Duc-Minh Luong, Van-Hai Duong, Xuan Bach Le Dinh, David Lo, Thang Huynh-Quyet.
  9. ICSE '22Toward the Analysis of Graph Neural Networks. Dat Nguyen, Thanh Le-Cong*, ThanhVu H. Nguyen, Xuan-Bach D. Le, Quyet-Thang Huynh.
  10. ICPR '20End-to-End Hierarchical Relation Extraction for Generic Form Understanding. Tuan-Anh Nguyen Dang, Duc Thanh Hoang, Quang Bach Tran, Chih-wei Pan, Dat Nguyen.
  11. MAPR '20PCA-based 3D Facial Reenactment From Single Image. Dat Nguyen, Tuan-Anh Nguyen Dang, Viet Sang Dinh.
  12. BMVC '19End-to-End Information Extraction by Character-Level Embedding and Multi-Stage Attentional UNet. Tuan Anh Nguyen Dang, Dat Nguyen.

Service

ICML 2026 Reviewer
NeurIPS 2026 Reviewer
AAAI 2026 PC Member / Reviewer