← All research
February 18, 20263 min read

Native ranks #1 on Spider 2.0

By Native Team

Native mini 1 is a research prototype still in development. It achieved 96.53% accuracy on Spider 2.0, placing #1 on the leaderboard and becoming the first system to break the 95% barrier.

Enterprises are betting on autonomous agents to run 24/7 across their data. That only works if they can trust the reasoning underneath. Native mini 1 is being built for exactly this.

Data Reasoning

BI tells you what happened. Data reasoning tells you why it happened, and what to do next, across ten different silos that don't talk to each other.

Most enterprises have spent months building agents that demo well but can't scale past a handful of medium-complexity tables. The underlying problem is the same. Their systems don't understand the data. Native mini 1 does.

We treat databases as a distinct modality. Building a deep understanding of raw data requires learning specific behaviours, relationships, and structures. Native mini 1 is the smallest in our family, at around 1/20th the cost of our larger configurations. It is the most accurate database reasoning system in the world.

Spider 2.0 results

To evaluate our infrastructure, we found the benchmark with the hardest questions, across the largest, lowest quality data. In other words, what actually exists inside enterprises. This was Spider 2.0.

Spider 2.0 tests a system's ability to reason over complex, messy, real-world database environments. Not clean academic datasets. The kind of sprawling, contradictory, multi-system data estates that enterprises actually operate on. Hundreds of tables, inconsistent schemas, missing labels, and questions that require multi-step reasoning across siloed sources.

Native mini 1 achieved 96.53% accuracy, placing #1 on the leaderboard and becoming the first team to break the 95% barrier. The benchmark has been attempted by research teams from ByteDance, Snowflake, Alibaba, Tencent, and others. Our larger reasoners solves every error free question in the benchmark but is reserved for enterprise deployments.

RankTeamAccuracy %
1Native mini 196.53
2Genloop v2 Pro95.06
3DAQUV + Gemini 3 Pro94.15
4Tencent with Contextual Scaling Engine93.97
5Paytm - Prism Swarm + Claude Sonnet 4.590.49
6AT&T & RelationalAI86.28
7ByteDance - ByteBrain Agent84.10
8Alibaba - AI Cheng Agent82.81
9Ant Group - LingXi Agent + Claude Sonnet 4.579.89
10Snowflake - Arctic-FLEX75.14

We are now working on other benchmarks as the gold answers from Spider 2.0 have now leaked into the training corpus of most LLMs.

Contact

Native mini 1 is a research prototype. Our enterprise model is significantly larger, built to work across every data system your organization runs on. Please email contact@usenative.ai if you'd like to learn more about our enterprise offering.