Blog
Longer-form technical notes on LLM research.
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Internalized Visual Thinking trains a multimodal LLM to predict next latent representations of future frames from unlabeled videos, then removes the image generation pathway at inference — keeping the foresight without paying for frame generation. A walkthrough of why explicit Visual CoT is expensive, what predictive supervision actually buys, and which design choices matter.
Read the note →