A Pareto frontier Database Reasoning model in 4B parameters
A research announcement
By Native
Our first submission, Native mini, with only 4B parameters, scored 76.86% on BIRD's hidden test set. That makes us the 3rd ranked team on BIRD's single-model leaderboard, behind Google and JPMorganChase, ahead of AWS, Tencent, Databricks, ByteDance, Snowflake and OpenAI, with a nearly 8× smaller model. This places Native on the Pareto frontier for model size and accuracy.
Enterprise AI is faster and more powerful when it can work on the data as it is, without waiting months for it to be cleaned or moved. That takes models that run 100% inside your environment, next to your data, without access to public models.
BIRD
BIRD has large messy databases, with complex ambiguous questions, covering a wide range of professional domains. Its test set is 100% hidden, and the BIRD team runs every submission themselves so its answers can't leak.
The Single Trained Model Track means you are only allowed one model answering on its own.
This is our first submission, with only a single 4B parameter model trained to be a pure data reasoning agent with no BIRD specific pipelines.
Our key innovation is using Reinforcement Learning to train our model to rely on understanding the data, not on understanding SQL. Its reasoning is grounded in the data, so it hallucinates less. BIRD's test databases are hidden, so we know it generalises to databases the model has never seen. We are excited to share more detail in future posts on how this is also working with larger models.
Demos
Because Native mini is only 4B parameters, it can run fast, at scale, and anywhere, all with high accuracy.
One agent at 2000 tok/s
The agent explores the database, checks its work against the data and answers, fast enough to keep up with your train of thought.
100 agents on one H100
A lot of data engineering comes down to looking at individual records. An agent that can do that, write SQL and reason across multiple tables, is incredibly useful, and with a 4B model you can run 100 of them at once.
In this airline industry example, someone wants to audit every single plane's July 2024 data in the US flights database for errors. Every single plane could have any kind of data error, so a single aggregated query / scoring metric won't work. It has to go look at all of the records.
The agent calls a tool to spin up 100 Native mini agents, one per plane. This all runs on a single H100. Each agent checks its plane's flight records with SQL, reasons over the data, and if required proposes fixes for each record. Running 100 agents at once gives us an average of 9500 tok/s per H100, and all 100 finish in about 10 seconds.
This works anywhere your data sits in a database and you need to go beyond summarising and actually look at every record to catch unique outliers and issues. Each agent writes its own SQL, pulls the records it needs and reasons over them. Say you are:
- A music streaming service that wants to reason over each listener's full listening history, and build a completely custom report or personalisation for every one of them
- A retailer mapping millions of products between your suppliers' tables and your own catalogue
- A bank or security team checking every account's transactions and every device's logs for fraud and new kinds of attacks
- A software company querying every user's events to see where they get stuck, and why
Offline on a phone
We don't imagine anyone using this on their phone (although it is pretty fun). The point is that these models run anywhere, completely owned by you. If it runs offline on a phone, it runs inside your environment, next to your data.
Native models run wherever your data is:
- In your own cloud account
- On servers in your own data centre
- In air-gapped networks with no internet
Native mini fits on a laptop or a phone, and our larger models still run on a single GPU.
What's next
We are building an incredible team across research, data engineering and software engineering. If you are obsessed with Data & AI, high agency and want real ownership, please email .
Native mini is a research model. For enterprise, if you would like to try our larger sovereign LLMs, RL-trained on your data, that are secure, incredibly fast, and accurate, please reach out at .