AlyronWork with us

Delta / evaluations & post-training

Measure what
your data teaches.

Delta turns enterprise work into tasks a model can attempt, answers a verifier can check, and evidence you can compare.

300private test tasks
643packages checked
12,359invalid mutations rejected

What we’ve measured

The first signal is already visible.

In our local procurement experiment, the model trained on real records solved all 20 development tasks. The matched synthetic-control model solved five. Delta makes that difference inspectable, down to the attempt and the verifier check.

Procurement pilot / recorded checkpoint

Real work. A measurable signal.

+75pp

Development-task pass-rate gap

Real-record training100.0% 20/20
Synthetic-control training25.0% 5/20
0%Whole-task success100%

Same 4B base model, 128 training cases per arm and matched supervised-token budgets. 20 development tasks scored by the programmatic verifier.

Development checkpoint after one matched continuation. The synthetic arm is still below the training-fit gate; the 300-task model test remains sealed.

The first enterprise workflow

Follow the records.
Resolve the differences.

Our procurement collection asks models to reconcile orders, receipt registers and purchase entries: calculate quantities, compare rates, flag inconsistencies and return the supporting record references.

Working state

Three records.
One reconciliation.

01Purchase orders
02Goods receipts
03Purchase entries
The model’s output
Recorded quantities↗
Rate comparisons↗
Exception flags↗
Evidence references↗
Programmatic verifierScore every attempt

Task-level evidence

Each result connects the starting state, submitted answer and grader checks.

A tested reward signal

Reference solutions and empty submissions were checked across 643 packages, alongside 12,359 invalid answer mutations.

A fixed comparison

Training arms share the base model, output contract and budget. A separate private exam is reserved for measuring what transfers.

Built to grow

One system.
More kinds of expert work.

The same loop connects data preparation, supervised training, reward-based learning and evaluation. We’re taking it from record reconciliation toward richer financial workflows, document coordination and controlled CAD edits.

01

Define the work

A goal, a reproducible starting state and an output contract.

02

Check the outcome

Executable rules grade the result and surface the failures.

03

Train and compare

Matched models learn from different data under the same recipe.

04

Measure transfer

Private tasks and external benchmarks test what the model can do next.

Build with Delta

Bring a capability
worth measuring.

For labs developing better models, and enterprises whose expert workflows can become the next training ground.