Reading models at the tensor level.
Architecture intelligence from ModelDNA — what config files, weight layouts, and modeling code reveal about frontier models, before anyone downloads a byte.
It Renamed DeepSeek's Architecture. Then the Weights Told the Truth.
SK Telecom's A.X K2 ships a novel-sounding architecture and a 'trained from scratch' claim — but the config is DeepSeek-V3.2 (they renamed DSA to 'SGA'). A per-tensor weight diff (median 113% Δ) shows the weights are genuinely its own: adoption, not a copy. The line only a weight diff can draw.
Who Actually Adopted MLA? We Counted.
We counted every fine-grained MoE model in the atlas. Original US adoption of MLA: zero (the 20% was all NVIDIA re-quants of Chinese models). MLA is a five-lab club — and even China isn't unanimous. A data-backed adoption survey.
What Solar-Open2 Reveals About the Frontier
A 250B model isn't chasing the frontier — which is exactly why its architecture is such a clean signal. Solar-Open2's choices reveal which frontier ideas have become table stakes (the MoE recipe) and which are still risky bets (MLA). Architecture as a revealed preference.
The Trick That Makes Kimi-K3’s 2.78 Trillion Parameters Almost Free
Everyone’s counting Kimi-K3’s parameters. We found the move that makes them cheap — a “latent MoE” that runs the experts in a compressed space — reverse-engineered from public metadata without downloading the model.