Blog21 August 2026

The future of AI is open.

Not as a licence footnote. As the default anyone serious will build on — and as the only reason a compact model we trained for frontend work could be dropped into live agent traffic, hold the quality bar of a frontier system, and process 7.8 billion tokens in an afternoon.

20invited~10,000users6 hourswindow7.8Btokens

Frontend is a bad place to hide. Layout, interaction, visual system, production code — you see the miss immediately. So that is the job we trained for. Reinforcement learning and a focused fine-tune on Qwen 3.8 — a compact open-weight model — then twenty keys into a coding-agent setup people already used. No lab harness. No cherry-picked prompts. If it was going to hold, it had to hold on their repos.

The keys moved. Within hours the endpoint was no longer a private trial. Roughly ten thousand people were on it, and we had to scale the serving plane while it was already hot. That is not the experiment we scheduled. It is the one we got.

In six hours the cluster processed 7.8 billion tokens — long agent loops, multi-file edits, the frontend work the checkpoint was tuned for. The feedback was unusually consistent: the model held. People who expected a toy open-weight run were shipping UI they would otherwise have reserved for a much larger closed model.

A compact, specialised model did the job. The giant one was not required.

You do not need a giant model to do giant work.

The useful unit of intelligence is not “the biggest model.” It is the right model for the job. A specialised open-weight checkpoint, trained for a narrow class of work, can match frontier-class quality on that class — at a fraction of the token cost of routing every turn through a general model.

That is only possible if the weights are open. You cannot reinforce, fine-tune, or deploy a closed model onto a workflow you actually own. The experiment exists because the weights could leave the lab and sit on our serving plane. Everything we mean by open-source AI follows from that fact.

What we mean by open

A model on your disk is a model nobody can deprecate, reprice overnight, or swap mid-quarter. That is the only version guarantee that has ever actually held — and it is why a run like this can exist at all. You cannot fine-tune a closed model onto your own workflow.

But open is more than the licence. Weights, harnesses, protocols, and data you can export whole. Each one is a door out of the building. A stack where all four are open is a stack you stay in because it is good, which is the only kind of retention worth having.

Choice is an engineering requirement, not a preference. Different work wants different models — a long reasoning step, a cheap loop, a coder, something small enough to sit next to regulated data. Standardising a company on one frontier model is a procurement outcome. It has never been an engineering one.

And control has to be structural. A promise about your data is worth as much as the architecture behind it. Either the prompts, the outputs and the logs live on hardware you control, or they do not — and no policy page changes which of those is true.

Lock-in does not arrive as a contract.

It arrives as a convenience. Nobody signs up for it — they adopt an agent that works, and the model, the cloud and the data path come welded to it. You adopted a tool. You also picked a model supplier, permanently, in the same click. Every prompt, every file, every log leaves through their pipe, under their retention schedule. Two quarters in, switching a model is a project. A switching cost that high is a tenancy.

The open stack is the same harness, but the model is a slot that takes any open weights. The ground is your own boundary. The exit is one line, not a rewrite.

What we are building

Open weights solved availability and none of the operational problems. You can download a capable model this afternoon and still have nowhere to put it: no serving that survives an agent looping for hours, no key custody, nothing a security review would sign. The choice on offer is a false one — take the closed stack and get the operations for free, or take the open models and build the floor yourself.

We are building the third option. Sideren serves enterprise teams who need privacy-first inference — on infrastructure we operate, or inside their VPC, dedicated host, or on-prem floor — with models deployed to their workflow the same way this checkpoint was trained for frontend. Quality where the task is specific. Token cost at a fraction of a frontier call. Data, keys and logs on hardware they control.

If the future of AI is going to be open, someone has to make specialised open models the easy option to run. That is the company.

More is coming soon. Follow Sideren on X for the latest updates.

— The Sideren team

Bring your own models. Keep your own ground.

Tell us the workflow and where inference has to live. We come back within 24 hours with a deployment plan.