Validation¶
The tests show that the pack's contracts hold on small fixtures. They do not show that the external tools produce correct biology. The validation program keeps these claims apart.
Evidence levels¶
| Level | Required evidence | Permitted claim |
|---|---|---|
| Fixture | Deterministic local inputs, schemas, command plans, normalized outputs and failure tests | The repository contract works on the fixture. |
| Scheduler integration | Pinned environment and databases, a completed scheduler job, exit state, elapsed time, peak RSS and checked outputs | The pinned external tools ran on the named scheduler profile. |
| Biological validation | Versioned truth labels, stratified scientific metrics, controls and documented limitations | The tested tool stack reached the reported metrics on the named truth set. |
One level does not imply the next. A completed Slurm job is not biological validation until its result is scored against truth labels.
What the tests cover¶
make test runs the skill, supplementary-doc and citation validators, the unit and fixture tests, and the routing benchmark. For skills that ship a driver or an artifact builder, the fixture tests check command plans, restart and reuse decisions, output normalization, schema validation and failure paths. They do not run the heavy external tools or their databases. A production run still needs pinned databases and containers, a scheduler job, and the QC gates in each skill.
Truth-set registry¶
validation/truth-sets.json records, for eight core skills, the candidate truth set, its evidence tier, biological strata, metrics, limitations, source release, license and artifacts. The skills are read QC and mapping, assembly, gene calling and ncRNA detection, functional annotation, phylogenomics, protein clustering and pangenomes, viromics, and interdomain horizontal gene transfer.
A truth set moves from candidate to ready only when every required artifact has an immutable URL and a locally verified SHA-256. Upstream MD5 values stay as provenance; they do not replace the SHA-256 check. All eight truth sets are candidate.
| Skill | Current surface | Needed for scheduler validation |
|---|---|---|
bio-reads-qc-mapping |
External-tool driver | Run the driver, score retained reads and mapping against truth, then test reuse. |
bio-assembly-qc |
External-tool driver | Run the driver and score MetaQUAST metrics against a gold assembly. |
bio-gene-calling |
Restartable external-tool driver | Run each domain route and compare CDS, protein, tRNA and rRNA calls with truth records. |
bio-annotation |
Artifact builder | Add an upstream annotation adapter before scoring CAFA or curated labels. |
bio-phylogenomics |
External-tool driver | Run marker trees and compare supported splits with the reference tree. |
bio-protein-clustering-pangenome |
Artifact builder | Run an orthology tool, then submit its predictions to QfO-compatible scoring. |
bio-viromics |
Artifact builder | Run geNomad and CheckV, then score labeled contigs before building the evidence bundle. |
bio-interdomain-hgt |
Artifact builder | Run the homology, context and tree stages; score simulations apart from curated empirical controls. |
Slurm jobs¶
A job manifest follows validation/schemas/slurm-job.schema.json and names:
- the validation, driver and truth set
- the cluster, account, partition, QOS, CPUs, memory and time
- a checksummed Pixi lock or container, and checksummed databases
- version commands, the analysis command and the expected outputs
Render the job without submitting it:
validation/scripts/submit_slurm_job.sh --dry-run \
validation/jobs/<ready-job>.json \
tasks/biological-validation/runs/<validation-id>/job.sbatch
The renderer rejects draft jobs and unresolved placeholders. Review the rendered script before you submit it with the same arguments and --submit:
--submitrenders the manifest again. When the script path already holds the reviewed script, the two must match byte for byte; when it does not, the new render is written and submitted unreviewed, so always run--dry-runfirst.--submitalso requiresOMICS_VALIDATION_SUBMIT_APPROVED=1. Set it only after the rendered script is approved.- The wrapper calls
sbatchwithout-M, so the job runs on the cluster of the host you submit from. Submit from a login node of the cluster named inscheduler.cluster, the one that holds the data, never from another cluster.
After the job ends, collect its evidence. --fetch-sacct queries sacct -M with the manifest's scheduler.cluster, because job IDs repeat across clusters.
uv run --script validation/scripts/collect_slurm_evidence.py \
validation/jobs/<ready-job>.json \
--job-id "$JOB_ID" --fetch-sacct \
--output tasks/biological-validation/runs/<validation-id>/run-evidence.json
The run record stores the cluster, Slurm state, exit code, elapsed seconds, peak RSS, requested resources, nodes, and output sizes and SHA-256 values. Scientific metrics are added only after these scheduler and artifact checks pass.
Current pilot¶
The first pilot, validation/jobs/phylogenomics-qfo-pilot.draft.json, stays draft until these values are known on a scheduler login node:
- the cluster, and a small-job account, partition and QOS on it
- the remote checkout and data paths
- the SHA-256 values of the QfO subset and the reference tree
- the SHA-256 of the solved Pixi lock
Phylogenomics goes first because it needs no large reference database.
Comparative-discovery thresholds need separate prokaryote, eukaryote, phage and Nucleocytoviricota campaigns.