CV

Education

Experience

  • 2026 - present

    Member of Technical Staff
    Decagon
    • Enhanced an opaque simulation framework for CX agents through LLM-based diagnosis of chat logs against a data-driven failure taxonomy, iterating with real users to create truly useful automations that accelerated debugging by 3x across 100+ deals.
    • Integrated customer telemetry into Decagon’s observability platform, unblocking agent hill-climbing over customer-indexed success metrics, which uncovered improvement opportunities that closed multi-million dollar contracts.
    • Consolidated alerting into a standard framework that eliminates coverage gaps and reliably detects incidents.
    • Built a system to enhance agents by translating customer test specs into formal simulations, identifying misalignment between programmed and anticipated agent behavior, and iteratively co-optimizing agents and simulations.
  • 2026 - present

    Co-founder and Lead Researcher
    Trace AI Labs
    • Created PACT, a multi-metric evaluation framework to assess whether enterprise AI assistants follow compliance rules in 48 realistic scenarios across regulated domains like healthcare, revealing no existing model is reliably safe (preprint under review).
    • Analyzed the psychology of why AI agents violate legal constraints under user pressure tactics, showing alignment between LLM behavior and prevailing human behavioral theories like deterrence and legitimacy (AIES 2026).
  • 2024 - 2026

    Atlanta, GA

    Research Assistant
    Georgia Tech — Human-Centered AI Lab
    • Advised by Dr. Mark Riedl.
    • Instrumented an explainable system that profiles LLMs’ strengths and weaknesses and selects the best model for tasks under cost constraints (MLSys YPS 2025), and generalized it to routing complex agent workflows (HCXAI at CHI 2026).
    • Traced where social reasoning in LLMs originates in pretraining corpora using training-data attribution (COLM 2026).
    • Developed counterfactual explanation methods that translate complex agentic-workflow behavior and errors into clear, actionable insights for developers through a visual, interactive interface (HCXAI at CHI 2026).
    • Led FLaME, the first holistic benchmark of LLMs over diverse financial analysis and reasoning tasks (ACL Findings 2025).
  • 2024 - 2025

    Atlanta, GA

    Research Assistant
    Georgia Tech Research Institute (GTRI)
    • Built an agentic sandbox with a chat interface for creating, executing, and debugging scientific experiments, infusing context from retrieved research papers and code to empower researchers to rapidly iterate on their ideas.
    • Proposed novel graph algorithms for decomposing scientific claims into atomic parts for automatic verification.
    • Architected MethodBench, a framework for evaluating LLM competence in determining the solution to open-ended, but previously solved, scientific problems (under review; DARPA Contract HR001125C0302).
  • 2025 - 2025

    New York, NY

    Software Engineer Intern
    Two Sigma Investments
    • Designed metrics to provide traders actionable feedback on their stock ideas and market timing, surfacing suboptimal trading patterns and informing more profitable quantitative models distilled from human trading intuition.
    • Built an AWS-native ETL pipeline using Lambda and Step Functions to process metrics in Postgres, incorporating observability, testing, and documentation to ensure reliable and rapid insight delivery to quantitative researchers.

Publications

Activities

  • Reviewing / Program Committees: ACL Rolling Review ERC Member (2026), ACM/AAAI Conference on AI Ethics and Society PC Member (2026), EleutherAI Summer of Open AI Research Application Reviewer (2026)
  • Awards / Fellowships: Supervised Program for Alignment Research (2026); Georgia Tech Fintech Research Fellow (2025)
  • Research Presentations: AAAI/ACM Conference on AI Ethics and Society (2026), ACM CHI Workshop on Human-Centered Explainable AI (2026), MLSys Young Professionals Symposium (2025), COLM Workshop on Agent Behavior (2026), COLM Workshop on Scientific Understanding of Foundation Models (2026)

Skills

Programming: Python, R, Java, SQL, Bash, C++, HTML, JavaScript, TypeScript, Go
Technologies: Git, GitHub, Docker, AWS, Hugging Face, ReactJS, Flask, REST APIs, Postgres, Coding Agents
Data & ML: PyTorch, TensorFlow, Scikit, Pandas, NumPy, Jupyter, MLFlow, Matplotlib, Seaborn, R-Studio
Natural Language: LLM and Agent Evaluation, LLM-as-judge, Fine-tuning, Agent Orchestration, Explainable AI