Multimodal Flow Modeling Just Cleared 82.8 on 150B Tokens
MF-1 posted 82.8, and multimodal flow modeling just moved out of my "huh, interesting" folder and into my architecture notes. One model, trained on 150B pretraining tokens, generating text and images under a single training objective. If your current stack is a language model wired to an