Layer 02 · Data: see the whole stackRunix Data · Early access

Domain data, cleaned and verified

Runix Data cleans, structures and builds training and evaluation data in six domains: code, finance, cybersecurity, legal, embodied AI and AI for Science, with code as the focus. Every record ships with its source, its licence and the checks it passed.

Built on Runix Pipeline, the data tooling its engagements run on.

example record
// one accepted record, pretty-printed
{
  "id": "code-task-000142",
  "source": {
    "repo": "example-org/example-parser",
    "base_commit": "3f9c2e1",
    "licence": "MIT"
  },
  "dedup": { "cluster": "c-5521", "kept": "1 of 3" },
  "checks": {
    "target_tests_with_patch": "pass",
    "target_tests_without_patch": "fail",
    "other_tests_with_patch": "pass",
    "secrets_scan": "clean",
    "statement_matches_tests": true
  },
  "verdict": "accepted",
  "dropped_reason": null
}

// and one that did not make it
{ "id": "code-task-000143", "verdict": "dropped",
  "dropped_reason": "target tests pass without the patch" }

An illustrative record. The repository, commit and IDs are placeholders.

Six domains Code Finance Cybersecurity Legal Embodied AI AI for Science

Six domains, six definitions of clean

The stages are the same everywhere: ingest, clean, structure, mask, report. What counts as a duplicate, a valid record or a leak changes with the field, and that is where each domain gets its own rules.

Repositories Focus

Code

Repository and task data for code models and coding agents, from training examples to evaluation tasks that have to run.

  • Licence and provenance recorded per file, and code whose licence does not allow the use left out
  • Secrets, credentials and personal data scrubbed before anything leaves the pipeline
  • Deduplication across forks, vendored dependencies and generated or minified code
  • Task statements that describe what the tests actually check

Training data · evaluation tasks

Code: rules and public references →

Code · verification

Every task is run, not read

A coding task only teaches or measures something if its tests can tell a right answer from a wrong one. Each task is executed in a clean container: the reference solution must make its target tests pass, the untouched repository must make those tests fail, and the rest of the suite must still pass.

A task whose target tests pass without its patch is dropped, and the report says so

Filings · transactions

Finance

Filings, statements, market data and transaction records, for models that have to read numbers as carefully as an analyst does.

  • Units, currencies and fiscal periods normalised before any two figures are compared
  • Restated figures kept apart from the originals they replace
  • One entity across tickers, legal names and registry identifiers
  • Account numbers and personal data masked, failing closed
Finance: rules and public references →

Advisories · logs

Cybersecurity

Advisories, vulnerability records, logs and threat reports, for models that triage, detect and explain.

  • The same vulnerability, reported by several sources, merged into one record
  • Affected version ranges normalised so they can be compared and queried
  • Indicators defanged, so no live payload or working link ships in a dataset
  • Labels for detection tasks, each with the rule that assigned it
Cybersecurity: rules and public references →

Trajectories · sensors

Embodied AI

Robot trajectories, teleoperation sessions and multi-sensor recordings, for policies that learn from demonstration.

  • Cameras, joint states and force readings aligned on one clock
  • Calibration kept with every episode, not in a separate file nobody ships
  • Long recordings segmented into episodes, with failed or unsafe ones filtered out
  • Action spaces normalised across robots and controllers
Embodied AI: rules and public references →

Antibodies · proteins

AI for Science

Antibody and protein records, where a bad merge changes the answer rather than the formatting. Our work here is limited to biological data.

  • Sequence formats converted and validated, not just parsed
  • One numbering scheme applied across sources that use competing ones
  • Accession numbers reconciled where two databases disagree
  • Names, aliases and identifiers normalised to one entity

We teach this work in a free course, CC BY-SA 4.0, taught in Chinese. Read the English overview

AI for Science: rules and public references →

Not listed

Another domain

The stages carry over from one field to the next; the rules do not. Tell us what the data is and what the model has to do with it, and we will say plainly whether we have the judgement for it.

Describe your data →

Domain judgement on shared tooling

Runix Pipeline runs the stages the same way on every engagement. Runix Data is the layer of judgement on top of it: the rules that decide, field by field, what survives and what is dropped.

You receive

Training dataFiles, a database or an endpointEvaluation dataKept apart from trainingQuality reportWith every delivery

Runix Data

CodeFinanceCybersecurityLegalEmbodied AIAI for Science

Runix Pipeline

IngestClean & dedupeStructureMaskReportDeliver

Your sources

DocumentsPDF, HTML, scansDatabasesExports and dumpsRepositoriesCode and historyRecordingsSensors and video
Read it from the bottom up. Raw material goes in through Runix Pipeline, the domain rules of Runix Data decide what survives, and what reaches you carries its provenance and a report.

What every delivery carries

The data is half of a delivery. The other half is what lets you check it without taking our word for it.

Provenance per record

Where each record came from, when it was collected and what was done to it, so a bad answer downstream can be traced back to the record that caused it.

A licence per record

The licence or permission each record was used under. Material whose terms do not allow your use is left out, not flagged and shipped anyway.

A quality report

Coverage, duplication, extraction confidence and the checks each record passed, plus what was dropped and why.

Personal data masked

Identified and masked before it reaches a training set. Detection fails closed: a record we cannot clear is held back, not passed through.

Evaluation kept apart

Evaluation data is split from training data by source rather than by row, so near-duplicates cannot sit on both sides of the split and inflate a score.

The schema, written down

Every field defined and every known gap stated, so the next team to use the data does not have to reverse-engineer it.

How early access works

Three steps, each with a person on the other end: Runix Data is scoped per engagement, not bought off a shelf.

01Send a sample and the task

A slice of the real data, not a description of it, and what the model has to do with the result: train on it, be evaluated on it, or both. We reply within one business day.

02Get a scoped plan and a quote

The rules we would apply, the checks each record has to pass, what we think is not worth doing, and the price, all agreed before any work starts.

03Receive the data and the report

Batch or continuous, as files, a database or an endpoint, whichever your training and evaluation jobs actually consume. Every delivery comes with its quality report.

Common questions

Which domains do you work in?

Code, finance, cybersecurity, legal, embodied AI and AI for Science, where our work is limited to biological data: antibody and protein records. Code is our focus. If your field is not on the list, tell us what the data is and we will say plainly whether we can do it well.

Do you clean our data, or build new data?

Both. We clean and structure data you provide or have the rights to use, and we build task data, such as verified coding tasks, to a specification agreed with you in writing before work starts.

What happens to the data we send?

It is processed only to do the work you asked for. It is not used to train models, ours or anyone else's, and it is not sold. The details are in our Privacy Policy and on our Security page.

How is it priced?

Per project or by volume, quoted before any work starts. You see the scope and the number together, so there is nothing to reconcile afterwards.

How does Runix Data relate to Runix Pipeline?

Runix Pipeline is the tooling underneath: the stages every engagement runs through, from ingest to delivery. Runix Data is the service built on it, with the domain rules and judgement calls that differ from one field to the next.

Send us a sample of the hard part

A slice of the real data and what the model has to do with it. We reply within one business day, and the scoped plan that follows includes the parts we think are not worth doing.

Request early access