Jun 2026 – Present
AI Safety Research Fellow
Second Look Research — University of Chicago Existential Risk Laboratory (XLab) · Chicago, IL
- Producing open-source replications of empirical AI safety research as part of a 10-week fellowship focused on strengthening the epistemic foundations of the field.
Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol LessWrong, August 2026
We replicate experiments from Greenblatt 2025 and Greenblatt 2026 on one baseline model from the original post, Opus 4.5. Our evaluations agree with the trends and quantitative values described in the original posts. We run similar evaluations on Claude Fable 5, Opus 5, and GPT-5.6-Sol and find that the newer models show a substantial jump in performance on some evals: Fable 5 gets 87.6% accuracy on Gen-Arithmetic with 10 problem repeats whereas previous SOTA was around 60%, and GPT-5.6-Sol experiences significant uplift from filler tokens and problem repeats on all four datasets, with filler tokens and repeats doubling performance from baseline on 3-hop.