Back to News Feed
TechCrunch AI17d agoAnna Heim

Kog is going deeper to squeeze more inference out of GPUs

While the tech industry has been captivated by the rise of purpose-built AI hardware—exemplified by the high-profile IPO of Cerebras earlier this year—a French startup named Kog is challenging the status quo. The company is betting that the industry’s existing fleet of datacenter GPUs, such as the Nvidia H200 and AMD MI300X, still holds vast, untapped potential for AI inference.

By focusing on deep-level software optimization rather than hardware manufacturing, Kog aims to prove that the perceived limitations of standard GPUs in handling agentic workflows are largely a misconception.

Unlocking Performance on Existing Hardware

In May, Kog made waves on Hacker News with a technical preview demonstrating that ultra-fast, single-request decoding is entirely achievable on the infrastructure enterprises already possess. While some observers were disappointed that the optimization didn't extend to consumer-grade laptop GPUs, the enterprise sector took immediate notice.

For businesses, inference speed and cost have become the primary bottlenecks in scaling AI. Kog’s promise to extract significantly higher performance from current hardware has already generated substantial interest.

"We had 200 tangible business leads," CEO Gaël Delalleau told reporters, noting that software engineering workflows are currently the most promising use case.

Many professional developers using tools like Claude Code are accustomed to waiting hours for complex tasks to complete. Anthropic has even introduced a "Fast Mode" at a premium price point to address these latency frustrations. Kog is positioning itself as the solution for these power users, as well as for emerging design platforms where generating apps or games from a prompt requires rapid, real-time outcomes to remain commercially viable.

The Pivot to Larger Models

Kog initially observed that while there is clear demand for speed, many prospective customers are hesitant to engage in the complex process of fine-tuning smaller models. In response, the startup has shifted its strategic focus.

  • Strategic Focus: Accelerating the development and deployment of larger, more capable models.
  • The Benchmark: The company’s initial demo achieved an impressive 3,000 tokens per second (TPS), though this was accomplished using the Laneformer 2B, a purpose-built model with 2 billion parameters.
  • The Goal: Proving that this high-speed methodology can scale to the massive LLMs that currently dominate the market.

Delalleau remains undeterred by skeptics who argue that large models are too cumbersome for standard GPU architectures. He maintains that modern GPUs possess immense memory bandwidth that is currently being underutilized by standard software stacks.

A "Hacker" Mindset Applied to Physics

Kog’s approach is deeply rooted in the unique background of its founder. Delalleau, who studied solid-state physics at École Polytechnique, spent years in the world of offensive cybersecurity. As a four-time finalist at the DEFCON CTF tournament, he developed a methodology that blends the rigor of physical science with the creative destruction of white-hat hacking.

The Kog Methodology:

1. Physics-Based Understanding: Analyzing the fundamental laws of the GPU hardware to maximize efficiency. 2. Low-Level Reverse Engineering: Dissecting hardware down to the assembly and binary levels to force the silicon to perform tasks it wasn't explicitly designed to handle.

This hands-on, granular approach is both a strength and a constraint. Because the team must spend weeks or even months "digging into the details" for every new GPU architecture, the startup’s current capacity is limited by its team of 11. However, Delalleau plans to eventually feed this methodology into agent-based pipelines, which would allow the company to support a wider array of chips and models simultaneously.

Sovereignty and the Road Ahead

Kog’s mission aligns with broader European efforts to establish technological sovereignty in the AI sector. With backing from Bpifrance, the French Tech 2030 program, and infrastructure support from Scaleway, the startup is well-positioned to navigate the competitive landscape.

While companies like ZML are also working on hardware-agnostic software to bypass Nvidia’s CUDA, Kog distinguishes itself through its hyper-focused, deep-level engineering approach. The next few months will be critical for the company’s growth trajectory.

Key Milestones for 2025:

  • September Target: Implementation of the first major LLM running at 10x speed.
  • Validation: Demonstrating tangible customer traction using this high-speed inference.
  • Series A: Utilizing these performance benchmarks to secure the next round of venture funding.

As the AI industry grapples with the high costs of inference, Kog’s bet on software-driven optimization could prove to be a vital bridge for enterprises looking to maximize their existing investments. If the team can successfully replicate their 3,000 TPS performance on larger, mainstream LLMs, they may well redefine the limits of what standard datacenter GPUs can achieve.