PhD student at University of British Columbia · Affiliated with Vector Institute · Research in AI Agents, LLMs, and RL
ClawBench: Can AI Agents Complete Everyday Online Tasks?
EMNLP 2026 Findings · V1: 153 tasks · V2: 130 tasks · 144 live websites · 15 categories
VidGround: Watch Before You Answer
Visually grounded post-training for video LLMs.
Dr. Claw: Your AI Research Assistant
EMNLP 2026 System Demonstrations · A full-stack research workspace for taking projects from idea to paper.
OpenSkill: Open-World Self-Evolution for LLM Agents
EMNLP 2026 · Builds both skills and verification signals from scratch, without target-task supervision.
RewardHarness: Self-Evolving Agentic Post-Training
COLM 2026 · A self-evolving agentic reward framework for image-editing evaluation.
Project Page · Paper · HF Paper · Releases
| Research question | Try the teaching resource | Original work |
|---|---|---|
| Does valid structured output preserve the requested fields and values? | StructEval CPU mini-lab · Two-minute walkthrough | StructEval |
| Does a retrieved paper actually support the claim being written? | Citation evidence worksheet | ScholarCopilot |
| What does a text-only video-question probe establish? | VidGround probe-log mini-lab | Watch Before You Answer |
These coauthor-maintained resources use handwritten examples, not new model experiments. The citation worksheet is a separate reading exercise, not a ScholarCopilot capability. Retaining a question after a text-only probe does not by itself prove visual dependence. Scripts, inputs, scope notes and original sources are linked from each resource; the video uses synthetic narration.
All publications and citation materials · Research program
- 2026.08: Five papers accepted at EMNLP 2026: WebWorld, OpenSkill, and VGI-BENCH; ClawBench in Findings; and Dr. Claw in System Demonstrations.
- 2026.04: New paper: ClawBench: Can AI Agents Complete Everyday Online Tasks?; 153 real-world tasks, 144 live websites, 7 frontier models. Best model: 33.3%.
- 2026.04: New paper: VidGround: Watch Before You Answer; visually grounded post-training for video LLMs.




