A Paris startup says it pushed 3,000 tokens per second through AMD MI300X and Nvidia H200 chips using software alone, with no custom silicon. The May demo exploded on Hacker News and, according to founder Gaël Delalleau, produced 200 tangible business leads.
Kog promises up to 30x faster LLM decoding on hardware enterprises already own. Delalleau, who studied solid-state physics before working in offensive security, likens the effort to Stanford’s Hazy Research lab, but focused deeper on GPU acceleration.
The pitch resonates where speed is money. Developers who wait hours for Claude Code results, and the premium Anthropic charges for fast mode, show that inference latency has a price. Prompt-to-app design partners also want quicker turns.
Early customers balked at fine-tuning small models, so Kog shifted its software push toward larger ones and open-sourced Laneformer, the 2B-parameter model behind the demo.
Its seed round was co-led by Paris VC Varsity VC. The founder’s argument: newer GPUs carry memory bandwidth nobody is using, and software can finally unlock it.