Since 2024, Korean companies and solo operators have followed roughly the same process when picking an AI tool: compare coding benchmarks across the major models, check processing speed and context window size, and test the output quality gap between free and paid tiers firsthand. There's good reason this process feels rational. Higher-performing models really do produce more useful results. Models with higher scores make fewer errors, summarize documents more accurately, and generate code more precisely. That belief is well-founded. In many practical domains, it still holds true today.

But over the past few months, something has happened that makes it worth questioning whether this standard explains the whole decision.

A Market Beyond the Benchmark

In July 2026, the American AI startup Cognition acquired a service called Poke. The deal was reportedly worth more than $100 million (roughly 138 billion won). Cognition is the company behind Devin, the software development agent. Devin writes code. It reviews code, runs tests, finds bugs, and fixes them.

Poke doesn't write code. It offered exactly one thing: "an AI that texts like a friend." Short replies, a light tone, a conversation that keeps flowing without dropping off. What Cognition paid 138 billion won for was the way that experience was designed.

Some read this deal as a simple acqui-hire. Using M&A to quickly bring in strong engineers is a familiar Silicon Valley pattern. But if that explanation were sufficient, it wouldn't account for why this particular sum went to this particular service. It makes more sense to read the Poke team's real strength as interaction design, not coding ability — that reading fits the price tag better.

Cognition's own public statements point the same way. The company said the acquisition wasn't meant to boost Devin's coding performance, but to change how Devin talks with developers. It has started treating "what an AI agent can do" and "what experience the person using it has" as two separate problems to solve.

Why Hundreds of Billions of Won Rode on How It Talks

Performance convergence.

Foundation models' coding performance has climbed fast over the past two years. According to OpenAI's GPT-4 technical report, published in March 2023, GPT-4 scored 67% on HumanEval. The same benchmark score kept rising across the major models released afterward. As of 2025, the HumanEval gap among top-tier models has narrowed considerably compared with a few years ago. That doesn't mean absolute performance no longer matters. It just means that, at the top end, deciding on a tool by a single number has gotten harder than it used to be.

A similar shift has been observed in other industries. Once washing performance leveled off across major appliance makers, vibration noise levels and how intuitive the control panel felt started carrying more weight in purchase decisions. Once the gap in refrigerator compressor efficiency narrowed, internal drawer layout and lighting design became what set products apart. Once functional performance becomes table stakes, competition shifts from features to the experience layer.

Cognition concluded that this same shift is now underway in the AI tools market. That Devin can write code is already proven. The next problem to solve was whether developers actually want to keep working with Devin. It acquired Poke to answer that question.

Behind this judgment is also a churn pattern in the AI tools market. For many AI services, there's a wide gap between initial adoption and retention three months later. Look closely at why people stop using an AI tool, and it's often not a lack of performance. They churn because the tool talks in a way that wears them down, its responses feel too formal, or it keeps losing track of their context. Performance was adequate — people simply didn't want to keep working with it.

So Are Performance Comparison Charts Now Useless?

No. The number of criteria simply went from one to two.

Benchmarks are still valid. Performance scores remain the most important criterion for batch pipelines that process large volumes of work at once, or for domains where output accuracy translates directly into cost. In areas like cost estimation, contract summarization, or code error detection — where checking every output by hand is impractical — the differences between models are real. In these domains, a correct answer matters more than a pleasant interaction.

But once performance clears a baseline threshold, repeat usage is what determines real-world value. A tool you use three times and abandon, even if it's slightly ahead on performance, creates less cumulative value than a comparably capable tool you can keep opening every day without friction. Performance means nothing if you stop opening the app.

The problem is that most Korean practitioners have no way to measure this second criterion. Benchmarks can be compared in a table. But "will I still want to open this tool in 30 days" is something you only find out by actually using it for 30 days. Because that's hard to measure, people default to what's measurable — the decision criteria end up shaped by ease of measurement rather than by what actually matters.

In environments where an IT department or procurement team assigns tools in bulk, this second criterion rarely makes it into the decision. The interaction experience an individual actually feels never reaches the final selection stage. But for solo operators, one-person PMs, and freelance planners who choose their own tools, this criterion can be added to the checklist right now.

As AI adoption accelerates, whatever can be automated keeps shifting to machines. Within that shift, a growing share of what people do is restructured into work performed jointly with AI tools. Which tool you picked is likely to matter less than how closely you can work with it — that's what will end up separating the quality of the output. As the territory machines handle expands, how a person contributes narrows down to one question: how well can you work with the machine.

Whether Cognition's bet pays off is still unknown. For interaction design to become a purchasing criterion as important as performance, the authority to choose a tool has to be sufficiently delegated to the end user. In environments where that delegation hasn't happened, the market Cognition is betting on doesn't yet exist at scale. This acquisition rests on the assumption that such delegation is happening fast.

The next time you evaluate an AI tool, don't stop at checking its performance score — actually use it for at least two weeks. Do you still want to open it after two weeks? Does talking to it feel like it's moving your work forward, or does it feel like just another task you have to handle? The answer to that question is just as real a basis for choosing a tool as the performance score.