645 / 2125

A startup claims it broke through a bottleneck that’s holding back LLMs

TL;DR

Miami-based Subquadratic came out of stealth in May claiming it has solved a mathematical efficiency bottleneck that has constrained LLMs for years. The issue is attention: in classic Transformers, compute rises sharply as context length grows. Subquadratic says it can make that scaling much cheaper. The evidence is still limited. MIT Technology Review says the startup has started showing technical receipts, but public details, benchmarks and independent replication remain thin.

Nauti's Take

If Subquadratic is right, context length stops being mostly a credit-card and GPU problem. But until open benchmarks land, this is an infrastructure promise with huge leverage and an equally huge burden of proof.

Briefingshow

Long context is one of the expensive parts of modern LLM products: more documents, longer agent runs and bigger memory windows push compute costs up. A real breakthrough in attention would change the cost curve, not just make one model a bit faster. That is why the claim needs hard external scrutiny.

Sources