← Blog
Industry CommentaryAI DevArchitecture

What the Kimi K3 Tool-Calling Benchmark Doesn't Measure

Simon Willison ran Kimi K3 through the pelican benchmark for agentic tool-calling. The model performs well. What the benchmark tests is the model's side of the agentic equation — not the API surface those agents work against. That's the variable teams should be thinking about.

Live demo

See CleenUI running before you talk to us.

Configure a demo environment in about twenty seconds — theme, spacing, the modules you care about, sample data for your industry — then launch the full dashboard in your browser. Nothing to install, no signup.

  • 20 sec setup
  • 15 themes
  • 20 modules
  • 80 industries
Pick a theme to start with