French startup Kog is betting it can extract dramatically more performance from conventional GPUs, promising 30x faster large language model inference through software optimization alone. The company, which drew attention in May with a technical demonstration on standard datacenter hardware, believes it can unlock new capabilities on existing chips without requiring purpose-built inference processors. With AI inference speed and costs now presenting a critical bottleneck for enterprise users, Kog's approach has attracted commercial interest from companies frustrated by delays in AI-powered workflows.
The startup's May demonstration achieved 3,000 tokens per second per request on AMD MI300X and Nvidia H200 GPUs, though this performance came from a custom small model with only 2 billion parameters called Laneformer 2B, which is now open source. The technical preview generated 200 tangible business leads, according to CEO Gaël Delalleau. Early feedback indicated software engineering would likely be the first use case, targeting customers who currently wait hours for results from tools like Claude Code and are put off by Anthropic's price premium for Fast Mode. The company also has design partners building tools that let users generate games and applications with prompts, where faster outcomes would translate directly to increased revenue.
Delalleau told TechCrunch that Kog has focused on accelerating larger models since the launch, after learning that prospective customers weren't prepared to fine-tune smaller models despite the initial demo's impressive speed. The CEO said newer GPUs possess increasing memory bandwidth that "only begs to be unlocked," contradicting skeptics who view GPUs as poorly suited for decoding tasks. Delalleau compared Kog's approach to Stanford's Hazy Research lab but with "an even deeper-level focus on GPU acceleration," drawing on his background in solid-state physics and offensive cybersecurity to reverse-engineer chips at the assembly language and binary code level.
This hands-on methodology requires several weeks or even months of dedicated GPU engineering research for each new chip, limiting the 11-person team to supporting a small number of processors in the near term. The company plans to eventually feed its methodology into agent-based pipelines that will enable support for more chips and models. Kog has backing from France's Bpifrance, French Tech 2030's program, and Scaleway, positioning it to benefit from Europe's push to build domestic capability in AI infrastructure. Co-leading the startup's seed round was Varsity VC, whose partner Kamel Zeroual was Delalleau's co-founder at his first startup, TechCrunch50 2009 participant Stribe.
Kog expects to implement its first major model at 10x speed in September, which Delalleau said will be crucial for demonstrating customer traction and raising a Series A round. The company acknowledges the market isn't quite mature yet, but sees growing demand as enterprises increasingly rely on AI for professional workflows where delays create real economic friction. Proving the approach works on full-scale LLMs remains the central challenge, requiring Kog to bridge a substantial gap from its 2-billion-parameter demo to models whose size often strains inference hardware. The software optimization thesis faces competition not just from purpose-built inference chips like those from Cerebras, which enjoyed a warm IPO reception in May, but also from other optimization-focused players like France's ZML, though Kog maintains its low-level focus sets it apart. If enterprises prove willing to swap specialized silicon for optimized software running on the datacenter GPUs they already own, the inference economics of the AI industry could shift considerably.

