Design Arena creators raise $7.9 million to bring taste to AI models
The journey of Intelligence, the startup behind the viral platform Design Arena, began in the most humble of settings: a dorm room just weeks before the founders' 2025 graduation. A group of college friends set out to build an AI-powered game engine, only to hit a wall that has plagued the industry since its inception. While their models could generate functional code and assets, the resulting games lacked the "fun factor" that defines a successful product.
This realization led co-founder Grace Li and her team to a pivotal conclusion: there is no algorithmic substitute for human intuition. They began experimenting with ways to crowdsource authentic human feedback at scale, a project that eventually evolved into Design Arena. Today, the platform boasts 5.3 million users worldwide, serving as a critical feedback loop for the world’s most advanced AI labs.
Solving the "Taste" Bottleneck
As it turns out, the founders weren't the only ones struggling to quantify quality. Many frontier AI labs were desperate for scalable, high-quality user data to refine their generative models.
“It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li explains. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.”
To capitalize on this momentum, the company—operating under the name Intelligence—announced on Monday that it has secured $7.9 million in seed funding. The round was led by Index Ventures, with notable participation from Conviction (Sarah Guo and Mike Vernal), A*, and Valkyrie.
How Design Arena Works
For the average user, the platform functions much like a sophisticated model router. Users interact with a ChatGPT-style interface to input prompts for websites, images, and various other visual formats. The platform then presents a series of "A vs. B" comparisons, forcing users to rank outputs from best to worst.
While the consumer-facing side is intuitive, the true engine of the business lies in the enterprise sector. Participating AI labs utilize this stream of human preference data to fine-tune their media-generating models. Because users are generally agnostic about which model produced a specific output—focusing purely on the quality of the result—the data provides an unfiltered look at what human beings actually prefer.
The platform’s success is reflected in its bottom line: Intelligence is currently generating $60 million in Annual Recurring Revenue (ARR), cementing its status as a vital infrastructure player in the AI ecosystem.
Beyond Automated Benchmarks
The data collected by Intelligence offers a unique advantage over traditional automated benchmarks. By requiring users to log in, the company can track how aesthetic preferences shift across different geographies and timeframes. For instance, Li notes that web dashboards in Asia often lean toward a more maximalist design aesthetic, a nuance that automated systems might overlook.
These human-led insights serve as a necessary safeguard against the risks of automated testing, which are increasingly susceptible to being "gamed" or manipulated—a vulnerability highlighted by the recent high-profile security breach at Hugging Face.
The Competitive Landscape
Despite the clear demand for human evaluation, the market remains volatile. The path to sustainability is not guaranteed; for example, Yupp shuttered its operations earlier this year despite raising $33 million from a16z crypto’s Chris Dixon and securing partnerships with frontier models.
However, the success of LM Arena, which raised $150 million in a Series A round just four months after launching its paid product, suggests that startups capable of effectively bridging the gap between raw AI output and human taste are poised to thrive. As Intelligence continues to scale, it remains a primary bellwether for the future of human-in-the-loop AI development.