Research at Alyron
Expert work.
New model capability.
We build environments from real operations, train models on the work, and measure what changes. Our aim is to make enterprise expertise a repeatable source of stronger AI.
A recorded starting point
Promising results.
A concrete next experiment.
Our first study has produced a clear development-stage gap. The research now focuses on qualifying both training arms and measuring the difference on the private exam.
Procurement pilot / recorded checkpoint
Real work. A measurable signal.
Development-task pass-rate gap
Same 4B base model, 128 training cases per arm and matched supervised-token budgets. 20 development tasks scored by the programmatic verifier.
Development checkpoint after one matched continuation. The synthetic arm is still below the training-fit gate; the 300-task model test remains sealed.
01 / Evaluation system
Delta
We’ve built the task, grading and comparison pipeline that lets us ask what a dataset teaches a model. Our first procurement study pairs real-record training with an independently originated synthetic control under the same rules and budget.
We’re working toward data selection based on measured learning value, with transfer tests that make improvements useful beyond the training collection.
Inside Delta ↗02 / Working research platform
The Delta Grounds
Our public platform brings together the benchmark library, task specifications and a runnable CAD development exam. The private arena grades attempts and exposes the checks behind each result.
The direction is a shared proving ground where labs can compare task collections, training recipes and model capability against fixed exams.
Explore the grounds ↗03 / Enterprise environments
A workflow becomes an experiment.
Our first private procurement benchmark has 300 test tasks, separated from training by supplier and document date. The tasks reconcile recorded quantities, rates and evidence across operational records. Across 643 packages, reference and no-op checks passed, and 12,359 invalid mutations were rejected.
We’re moving toward harder exceptions and longer workflows: the work that needs context across records, systems and decisions.
See the recorded results ↗04 / CAD & architecture
Preserve the intent.
Change the design.
Our live six-task CAD exam tests controlled parameter edits, feature preservation and analytic clearances, with per-criterion grading and OpenSCAD export. The CAD directory maps drawing understanding, parametric editing and structural modeling benchmarks.
We’re taking this toward native CAD execution, real revision histories, drawing coordination and simulation-backed checks.
Explore the CAD exam ↗05 / Transfer evaluation
The next task is the test.
We’ve prepared a 200-problem FinQA transfer suite with its official execution evaluator. Our benchmark directory also covers spreadsheet manipulation, workplace agents and CAD, so we can choose external tests that match the capability we’re training.
We want to show that learning from enterprise work improves performance on unfamiliar problems, and make those comparisons reproducible.
Browse the benchmark library ↗06 / SFT + reinforcement learning
A better attempt.
A stronger model.
Our local SFT runs have already produced models we can evaluate against their task-level outputs. The real-record arm reached 125/128 correct answers on its training tasks and 20/20 on development; the synthetic arm reached 113/128 and 5/20.
We’re extending that work into reinforcement learning, where models improve by attempting the workflow, receiving a checkable reward and trying again.
Explore the training comparison ↗Collaborate
What should AI
learn to do next?
Bring the workflow, the hard cases or the capability target. We’ll build the experiment around it.
