AI Research Engineer (Intern)
June 2026 – August 2026Stanford Tax Lab (STAX), Stanford Graduate School of Business · Advisor: Prof. Rebecca Lester
- Bay Area cities publish their council records on six different civic platforms, and roughly 4,900 meetings existed only as video, with no text at all. I pulled records from 82 cities across Legistar, CivicClerk, CivicPlus, Granicus, YouTube and Vimeo into one schema, then transcribed ~9,500 hours of audio with Whisper large-v3-turbo on the Stanford GSB Yen cluster (H200 and A40 GPUs). Final corpus: 68,000+ documents, 270M words, spanning 2000 to 2026, and in active use by a researcher in the lab working on local business tax sentiment.
- A retrieval system that only confirms what you already believe is worse than none. I built an agentic retrieval system over a corporate ownership graph covering 100M+ entities, 2007 to 2023 — a LangGraph architecture with a set of research tools — designed to disconfirm rather than confirm: it flags the provenance of each claim and surfaces ownership relationships that the source documents do not actually support. Ran under Stanford's enterprise data agreements, keeping restricted data on approved systems.
- Tax credit transfers get disclosed in prose, not in a table. I built scraping and LLM extraction over 10-K and 8-K filings that turns those unstructured disclosures into a structured dataset, surfacing 1,000+ tax credit transfer instances for a job-market paper.
- Set up the lab's GitHub organization, templated repos, and reproducibility standards, adopted across a research team with a wide range of technical backgrounds. Rebuilt the lab website on Astro and GitHub Pages and handed it off.
Not public — data licensing.