← Blog
Industry CommentaryAI DevArchitecture

What the Kimi K3 Tool-Calling Benchmark Doesn't Measure

Simon Willison ran Kimi K3 through the pelican benchmark for agentic tool-calling. The model performs well. What the benchmark tests is the model's side of the agentic equation — not the API surface those agents work against. That's the variable teams should be thinking about.