CV
Education
-
2023 - 2026 Atlanta, GA
Experience
-
2026 - present Member of Technical Staff
Decagon
- Enhanced an opaque simulation framework for CX agents through LLM-based diagnosis of chat logs against a data-driven failure taxonomy, iterating with real users to create truly useful automations that accelerated debugging by 3x across 100+ deals.
- Integrated customer telemetry into Decagon’s observability platform, unblocking agent hill-climbing over customer-indexed success metrics, which uncovered improvement opportunities that closed multi-million dollar contracts.
- Consolidated alerting into a standard framework that eliminates coverage gaps and reliably detects incidents.
- Built a system to enhance agents by translating customer test specs into formal simulations, identifying misalignment between programmed and anticipated agent behavior, and iteratively co-optimizing agents and simulations.
-
2026 - present Co-founder and Lead Researcher
Trace AI Labs
- Created PACT, a multi-metric evaluation framework to assess whether enterprise AI assistants follow compliance rules in 48 realistic scenarios across regulated domains like healthcare, revealing no existing model is reliably safe (preprint under review).
- Analyzed the psychology of why AI agents violate legal constraints under user pressure tactics, showing alignment between LLM behavior and prevailing human behavioral theories like deterrence and legitimacy (AIES 2026).
-
2024 - 2026 Atlanta, GA
Research Assistant
Georgia Tech — Human-Centered AI Lab
- Advised by Dr. Mark Riedl.
- Instrumented an explainable system that profiles LLMs’ strengths and weaknesses and selects the best model for tasks under cost constraints (MLSys YPS 2025), and generalized it to routing complex agent workflows (HCXAI at CHI 2026).
- Traced where social reasoning in LLMs originates in pretraining corpora using training-data attribution (COLM 2026).
- Developed counterfactual explanation methods that translate complex agentic-workflow behavior and errors into clear, actionable insights for developers through a visual, interactive interface (HCXAI at CHI 2026).
- Led FLaME, the first holistic benchmark of LLMs over diverse financial analysis and reasoning tasks (ACL Findings 2025).
-
2024 - 2025 Atlanta, GA
Research Assistant
Georgia Tech Research Institute (GTRI)
- Built an agentic sandbox with a chat interface for creating, executing, and debugging scientific experiments, infusing context from retrieved research papers and code to empower researchers to rapidly iterate on their ideas.
- Proposed novel graph algorithms for decomposing scientific claims into atomic parts for automatic verification.
- Architected MethodBench, a framework for evaluating LLM competence in determining the solution to open-ended, but previously solved, scientific problems (under review; DARPA Contract HR001125C0302).
-
2025 - 2025 New York, NY
Software Engineer Intern
Two Sigma Investments
- Designed metrics to provide traders actionable feedback on their stock ideas and market timing, surfacing suboptimal trading patterns and informing more profitable quantitative models distilled from human trading intuition.
- Built an AWS-native ETL pipeline using Lambda and Step Functions to process metrics in Postgres, incorporating observability, testing, and documentation to ensure reliable and rapid insight delivery to quantitative researchers.
Publications
-
2026 PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
Under review
Okamoto, M., & Erol, A. K.
-
2026 Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
AIES 2026 (also at the COLM Workshop on Agent Behavior)
Okamoto, M., Erol, A. K., & Erol, K.
-
2026 Where Does Social Reasoning Come From? Capability Provenance in Language Models
Conference on Language Modeling (COLM) 2026
Matlin, G., Chakraborty, C., Eom, S., Okamoto, M., et al.
-
2026 Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
COLM Workshop on Scientific Understanding of Foundation Models 2026
Okamoto, M., & Sarti, G.
-
2026 Explainable Model Routing for Agentic Workflows
HCXAI Workshop at CHI 2026
Okamoto, M., Erol, A., & Riedl, M.
-
2026 Method Bench: Evaluating Path-to-Feasibility Reasoning in the Sciences
Under review
Johnson, B., Okamoto, M., Kerce, J. C., & Fekri, F.
-
2026 Counterfactual Explanations for Agentic Workflows
HCXAI Workshop at CHI 2026
Singh†, M., Kim†, G., Okamoto†, M., et al. († equal contribution)
-
2026 DeepVerify: Evidence-Based Expert-Level Scientific Claim Verification
Preprint
Xiong, S., Gungordu, O., Johnson, B., Okamoto, M., Kerce, J. C., & Fekri, F.
-
2025 FLaME: Holistic Finance Language Model Evaluation
ACL Findings 2025
Matlin†, G., Okamoto†, M., Pardawala†, H., Yang, Y., & Chava, S. († equal contribution)
-
2025 Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
MLSys Young Professionals Symposium (YPS) 2025
Okamoto, M., Erol, A., & Matlin, G.
-
2025 Enhancing Military Family Readiness and Resilience Programs Through LLM-Generated Synthetic Data and Natural Language Interfaces
Military Health System Research Symposium (MHSRS) 2025
Iyer†, A., Okamoto†, M., et al. († equal contribution)
-
2025 Genome-Wide Association Study of Age-Related Hearing Loss in CFW Mice Identifies Multiple Genes and Loci, Including Prkag2
Journal of the Association for Research in Otolaryngology (JARO)
Polesskaya, O., Boussaty, E., Cheng, R., et al.
-
2023 RATTACA: Genetic Predictions in Heterogeneous Stock Rats Offer a New Tool for Genetic Correlation and Experimental Design
bioRxiv
Johnson†, B. B., Sanches†, T. M., Okamoto†, M., et al. († equal contribution)
Activities
- Reviewing / Program Committees: ACL Rolling Review ERC Member (2026), ACM/AAAI Conference on AI Ethics and Society PC Member (2026), EleutherAI Summer of Open AI Research Application Reviewer (2026)
- Awards / Fellowships: Supervised Program for Alignment Research (2026); Georgia Tech Fintech Research Fellow (2025)
- Research Presentations: AAAI/ACM Conference on AI Ethics and Society (2026), ACM CHI Workshop on Human-Centered Explainable AI (2026), MLSys Young Professionals Symposium (2025), COLM Workshop on Agent Behavior (2026), COLM Workshop on Scientific Understanding of Foundation Models (2026)
Skills
Programming: Python, R, Java, SQL, Bash, C++, HTML, JavaScript, TypeScript, Go
Technologies: Git, GitHub, Docker, AWS, Hugging Face, ReactJS, Flask, REST APIs, Postgres, Coding Agents
Data & ML: PyTorch, TensorFlow, Scikit, Pandas, NumPy, Jupyter, MLFlow, Matplotlib, Seaborn, R-Studio
Natural Language: LLM and Agent Evaluation, LLM-as-judge, Fine-tuning, Agent Orchestration, Explainable AI