Which annotations deserve a second look?

A group-safe framework that ranks potentially inconsistent annotations for expert review without changing source labels.

Annotation auditing is a way to decide what a human should inspect first. It is not a way to replace the human decision.

Large histopathology datasets contain many already segmented nucleus instances, each paired with a class label. Even careful annotation can include ambiguity, inconsistent conventions or ordinary data-entry noise. Reviewing every instance again is expensive, so the useful question is not “can a model declare a label wrong?” It is “can a model create a better review queue than random sampling?”

AANCA evaluates that question under controlled conditions. It intentionally changes a known subset of class labels, hides that intervention from the auditor, and then measures whether the changed labels move toward the front of a fixed review queue. Because the benchmark records exactly what it changed, retrieval can be scored without pretending that model disagreement is biological truth.

The pre-corruption label is an experimental reference label, not guaranteed biological truth. It is used only to define the controlled benchmark, simulated restoration and final evaluation. In a real audit, a high score means recommended for expert review. It never means “confirmed error.”

Study boundary. The PanNuke primary benchmark, a frozen NuCLS multi-rater evaluation, a frozen MoNuSAC controlled benchmark and a new-source PUMA controlled confirmation have been completed. PUMA supports transfer under controlled label noise; NuCLS does not support natural-error or downstream-improvement claims. Blinded natural-case expert review and prospective clinical workflow evaluation have not been performed. The PanNuke primary analysis remains exploratory because outcomes were exposed during technical recovery.

Benchmark setup

01Data
PanNuke

Verified official release; five positive nucleus classes across 19 tissue types.

02Unit
Already segmented nuclei

Class-label consistency only. Segmentation quality and diagnosis are outside scope.

03Primary model
Frozen ResNet-18 + logistic regression

ImageNet context embeddings with balanced multinomial fitting.

04Prediction design
Five-fold group-safe OOF

A scored nucleus and its whole source patch are absent from its training fold.

05Review budget
5% primary queue

Guided and random review receive the same integer budget.

How the audit works.

The first two states show three parts kept separate for every benchmark record: the pre-corruption reference label, the observed label exposed to the model and metadata describing the controlled intervention. When corruption is injected, the observed label may be replaced and the metadata records what changed; the reference label remains fixed. It defines the benchmark target, not guaranteed biological truth.

The centre shows one of five group-safe out-of-fold splits. Four neutral folds fit the model; the violet fold contains the source group being scored. Neither that nucleus nor any other nucleus from the same source patch appears in that model's training data. Cycling the held-out fold produces one probability vector for every audited instance. The primary risk score rises as the model assigns less probability to the observed label; a higher score means earlier review, not a confirmed error.

On the right, the guided queue orders instances by risk while the random queue provides a matched baseline. Both receive exactly the same integer review budget. The five squares are schematic: the primary queue contains 5% of eligible instances, not five nuclei. Retrieval is then compared against the injected changes known only to the evaluator.

The final fork separates retrieval from downstream utility. The magnifier represents how many injected changes the queue retrieves. In the PanNuke controlled restoration experiment, only reviewed injected changes are restored, all unreviewed observed labels remain unchanged and the downstream classifier is evaluated on the untouched, uncorrupted final reference fold. Natural-data auditing never performs this restoration: AANCA only recommends potentially inconsistent annotations for expert review and never changes source annotations automatically.

Conceptual workflow showing separate label states, four training folds scoring a held-out source group, equal guided and random review budgets, retrieval measurement and a separate downstream evaluation on an untouched final fold.
Conceptual workflow. Violet denotes the controlled intervention, held-out scoring state and guided queue; it does not indicate biological truth or an automatic correction.

Four quantities answer four different questions.

Separating them prevents a strong ranking result from being mistaken for clinical or downstream utility.

Average precision describes the complete ordering of the review queue. Recall at the primary 5% budget describes the operational slice an expert would actually inspect. The random-review baseline tells us whether that prioritisation is better than spending the same review effort without a model.

Restoration utility is deliberately separate. It asks what happens after simulated review when found injected changes are restored and the downstream classifier is retrained. A ranking can retrieve controlled inconsistencies efficiently while still failing to improve the later classifier.

Intervals and adjusted tests also answer different questions. The displayed 95% intervals come from paired whole-group bootstrap resampling; registered Holm-adjusted values are one-sided. At 0% corruption, average precision is not applicable because the controlled positive event does not exist.

On the original PanNuke benchmark, better triage did not improve the downstream model.

This is the most important negative PanNuke result. Ranking injected inconsistencies and improving a later classifier are related, but they are not equivalent objectives.

The H4 experiment restored exactly the same 5% review budget under audit-guided and random selection, then trained the same downstream model and evaluated it on the untouched final reference fold. Audit-guided restoration was not favoured.

Audit-guided minus random review
-0.002156

Audit-guided restoration produced macro-F1 0.524431, below the random-review mean of 0.526587.

Plain-language interpretation

The audit-guided result was 0.216 percentage points lower. The saved 95% interval remained below zero, so this registered test did not support downstream benefit.

Macro-F1 after equal-budget label restoration

0.523Macro-F10.539
Reference · uncorrupted0.537499
Baseline · corrupted labels0.526717
Random review · mean0.526587
Audit-guided review0.524431
Audit-guided minus random review-0.002156
-0.003200.0003

95% CI [-0.002859, -0.001393] · 100 random-review repetitions

This result does not erase the ranking evidence. It shows that retrieving more injected label changes near the front of a queue and improving final-fold classification are not interchangeable objectives. The first evaluates triage efficiency; the second depends on how the reviewed cases affect a later training distribution.

Only reviewed injected corruptions were restored; unreviewed observations remained unchanged, and all downstream conditions used the same representation, learner, seeds, review count and untouched final fold. The adverse direction is retained without selecting a post-result explanation.

The remaining registered questions

H1 through H7 were designed to separate retrieval, heterogeneity, corruption difficulty, restoration, score combination, representation availability and target highlighting. Reading them together prevents one favourable comparison from becoming the story of the entire study.

The pattern is mixed. H1 supports better-than-random prioritisation inside the controlled benchmark, H5 favours the fixed hybrid, and H3 indicates that some injected mechanisms were harder to retrieve. H4 is adverse, H7 is unresolved, and H6 is unavailable because no pathology encoder passed the frozen eligibility gates.

Each answer below states the conclusion, the exact saved evidence and the interpretation boundary. None of the positive ranking results establishes that a naturally occurring annotation is wrong.

What the study actually learned.

Can the audit score find injected label changes earlier than random review?

Yes, within this controlled benchmark. Self-confidence ranking beat random review in all 12 registered comparisons: average-precision gains ranged from 0.036206 to 0.301884, and every saved 95% bootstrap interval was above zero. This means injected changes were prioritised more efficiently; it does not prove that a natural annotation is wrong. Byte-identical instance-dependent seeds are not independent replications.

Where did ranking performance vary?

Across 32,760 reportable subgroup estimates, ranking performance varied with nucleus class, tissue, corruption mechanism and corruption rate. The practical lesson is that audit difficulty depends on context. These are descriptive summaries only: the study did not register an omnibus test and cannot interpret the variation as a causal biological effect.

Which injected changes were harder to retrieve?

Confusion-targeted and instance-dependent changes were harder to rank than symmetric changes, meaning their average precision was lower. All 6 registered point differences favoured symmetric corruption; five saved 95% intervals were above zero and one crossed zero. This describes the controlled benchmark and does not establish that one natural error type is universally harder.

Did a better review queue improve the later classifier?

No. The main result above reports the complete comparison. At the same 5% review budget, audit-guided restoration was lower by -0.002156 macro-F1, so better retrieval did not translate into better downstream classification here.

Did combining two ranking signals help?

Yes, for review prioritisation. The equal-weight combination of self-confidence and fold-safe nearest-neighbour disagreement outperformed self-confidence alone in all 12 registered comparisons. Average-precision gains ranged from 0.020983 to 0.065502, and every saved 95% bootstrap interval was above zero. This supports the fixed hybrid within this benchmark; it does not authorise automatic correction or diagnosis.

Could a pathology-specific encoder be evaluated fairly?

No result is available. None of the candidate pathology encoders passed every frozen access, licence, reproducibility, hardware and smoke-test gate, so all 3 registered H6 entries remain unavailable. They are not zero or negative results, and no substitute was selected after the other outcomes became visible.

Did highlighting the target nucleus add useful signal?

No clear benefit was detected. Across three seeds, the average-precision differences ranged from -0.005371 to 0.002997; all 3 saved 95% intervals crossed zero, and no Holm-adjusted one-sided p-value was below 0.05. In this benchmark, explicit highlighting did not improve ranking over the context representation.

Every preregistered comparison entry.

Each point is an average-precision difference; each line is its saved two-sided 95% percentile-bootstrap interval. Missing H6 points are shown as unavailable rather than as zero.

The atlas is the transition from narrative findings to detailed evidence. Rows remain grouped by hypothesis so that effect direction, interval width, unavailable cells and repeated seed structure can be inspected together instead of reduced to a single headline number.

The plot is intentionally bounded inside the article. Scrolling within it reveals all 36 registered entries while keeping the surrounding explanation in view; the complete numeric record appears later in the evidence table.

H1Point estimate and saved 95% bootstrap interval
Self-confidence vs random review · Confusion-targeted · Seed 4040.284387
Self-confidence vs random review · Confusion-targeted · Seed 4050.290995
Self-confidence vs random review · Confusion-targeted · Seed 4060.290707
Self-confidence vs random review · Group-conditional · Seed 4040.301884
Self-confidence vs random review · Group-conditional · Seed 4050.295644
Self-confidence vs random review · Group-conditional · Seed 4060.293129
Self-confidence vs random review · Instance-dependent · Seed 4040.036206
Self-confidence vs random review · Instance-dependent · Seed 4050.036206
Self-confidence vs random review · Instance-dependent · Seed 4060.036206
Self-confidence vs random review · Symmetric · Seed 4040.292982
Self-confidence vs random review · Symmetric · Seed 4050.298066
Self-confidence vs random review · Symmetric · Seed 4060.301792
H3Point estimate and saved 95% bootstrap interval
Symmetric vs confusion-targeted corruption · Seed 4040.008595
Symmetric vs confusion-targeted corruption · Seed 4050.007071
Symmetric vs confusion-targeted corruption · Seed 4060.011085
Symmetric vs instance-dependent corruption · Seed 4040.256775
Symmetric vs instance-dependent corruption · Seed 4050.261788
Symmetric vs instance-dependent corruption · Seed 4060.265437
H5Point estimate and saved 95% bootstrap interval
Fixed hybrid vs self-confidence · Confusion-targeted · Seed 4040.052548
Fixed hybrid vs self-confidence · Confusion-targeted · Seed 4050.054609
Fixed hybrid vs self-confidence · Confusion-targeted · Seed 4060.054104
Fixed hybrid vs self-confidence · Group-conditional · Seed 4040.060268
Fixed hybrid vs self-confidence · Group-conditional · Seed 4050.065135
Fixed hybrid vs self-confidence · Group-conditional · Seed 4060.065502
Fixed hybrid vs self-confidence · Instance-dependent · Seed 4040.020983
Fixed hybrid vs self-confidence · Instance-dependent · Seed 4050.020983
Fixed hybrid vs self-confidence · Instance-dependent · Seed 4060.020983
Fixed hybrid vs self-confidence · Symmetric · Seed 4040.062592
Fixed hybrid vs self-confidence · Symmetric · Seed 4050.064366
Fixed hybrid vs self-confidence · Symmetric · Seed 4060.064656
H6Point estimate and saved 95% bootstrap interval
Pathology encoder vs ImageNet · Seed 404Unavailable by frozen designUnavailable
Pathology encoder vs ImageNet · Seed 405Unavailable by frozen designUnavailable
Pathology encoder vs ImageNet · Seed 406Unavailable by frozen designUnavailable
H7Point estimate and saved 95% bootstrap interval
Target-highlighted vs context · Seed 4040.002997
Target-highlighted vs context · Seed 405-0.000004
Target-highlighted vs context · Seed 406-0.005371

Statistical reading. Registered Holm-adjusted p-values are one-sided, while the displayed 95% intervals are two-sided summaries. Those two summaries need not produce identical verbal labels.

Saved performance estimates varied across contexts.

The saved subgroup estimates span nucleus classes, tissues, corruption mechanisms and corruption rates. They describe where performance differed in this benchmark; they do not establish a causal biological dependency.

Subgroup average precision is reported only when the preregistered support rule is met: at least 100 samples and at least 10 injected corruptions. Otherwise the study retains counts without inventing an unstable estimate.

The ranges expose heterogeneity that would disappear inside one aggregate score. They are descriptive, not a biological ranking: the preregistration did not define an omnibus test that could support a causal explanation for tissue, class, mechanism or corruption-rate differences.

Observed average-precision ranges

0 APRange across saved subgroup estimates1 AP
Nucleus class6,300reported0.0136-0.8551
Tissue type23,940reported0.0185-0.7712
Corruption mechanism1,260reported0.0331-0.6429
Corruption rate1,260reported0.0331-0.6429

32,760 reportable estimates from 33,670 saved rows. These ranges are descriptive. They are not an omnibus test or a biological ranking.

Three registered seeds produced one realisation.

The three rows are retained for auditability, but the byte-identical files are not independent realisations and must not be interpreted as three replications.

One deterministic output, preserved through three registered paths.

The ranking and out-of-fold prediction files match byte for byte. The disclosure changes interpretation, not the stored record: these rows document one realisation rather than three replications.

Inspect exact seed identities and SHA-256 hashesThree registered paths · one byte-identical realisation
SeedCell IDRanking SHA-256OOF SHA-256
404primary_0187_53201213be7769766d68a24679b5…730459794225e338…
405primary_0193_afae7a24423869766d68a24679b5…730459794225e338…
406primary_0199_6029f5261b3869766d68a24679b5…730459794225e338…

Every displayed result remains inspectable.

Machine-readable summaries, exact identifiers and the complete 36-entry table remain available for audit. Package verification checks the published files; it does not recompute the primary study.

Every value in the article is read from the accepted sealed run and carried into evidence.json. The table below traces each statement to its hypothesis, raw identifier, seed, point difference, confidence interval, adjusted p-value and bootstrap count.

The timestamped July freeze records a null commit, a dirty tree and untracked files; the first public Git commit followed on 19 August 2026. Internal hashes preserve file identity, but the later public history is not independent evidence that outcomes were unseen.

The public primary evidence release contains all completed-cell OOF predictions and rankings, the full group bootstrap, subgroup table and H4 restoration arrays. A standalone evidence recalculator that does not import the primary analysis package recomputes the saved H1-H7 comparison statistics; this is not third-party validation. Fold checkpoints were not retained, and a second image-to-result execution still requires a lawful PanNuke copy.

Inspect the complete H1 / H3 / H5 / H6 / H7 table 33 reported · 3 unavailable
36 / 36 rows
HypothesisComparisonStatusΔ AP95% bootstrap CIHolm-adjusted pIterations
H1Self-confidence vs random review · Confusion-targeted · Seed 404h1_self_confidence_minus_random_confusion_seed_404Reported0.284387[0.270908, 0.296645]0.0059972,000
H1Self-confidence vs random review · Confusion-targeted · Seed 405h1_self_confidence_minus_random_confusion_seed_405Reported0.290995[0.278642, 0.302825]0.0059972,000
H1Self-confidence vs random review · Confusion-targeted · Seed 406h1_self_confidence_minus_random_confusion_seed_406Reported0.290707[0.277797, 0.303165]0.0059972,000
H1Self-confidence vs random review · Group-conditional · Seed 404h1_self_confidence_minus_random_group_conditional_seed_404Reported0.301884[0.288905, 0.314960]0.0059972,000
H1Self-confidence vs random review · Group-conditional · Seed 405h1_self_confidence_minus_random_group_conditional_seed_405Reported0.295644[0.283044, 0.308389]0.0059972,000
H1Self-confidence vs random review · Group-conditional · Seed 406h1_self_confidence_minus_random_group_conditional_seed_406Reported0.293129[0.280287, 0.305916]0.0059972,000
H1Self-confidence vs random review · Instance-dependent · Seed 404h1_self_confidence_minus_random_instance_dependent_seed_404Reported0.036206[0.031088, 0.041555]0.0059972,000
H1Self-confidence vs random review · Instance-dependent · Seed 405h1_self_confidence_minus_random_instance_dependent_seed_405Reported0.036206[0.031088, 0.041555]0.0059972,000
H1Self-confidence vs random review · Instance-dependent · Seed 406h1_self_confidence_minus_random_instance_dependent_seed_406Reported0.036206[0.031088, 0.041555]0.0059972,000
H1Self-confidence vs random review · Symmetric · Seed 404h1_self_confidence_minus_random_symmetric_seed_404Reported0.292982[0.280050, 0.305129]0.0059972,000
H1Self-confidence vs random review · Symmetric · Seed 405h1_self_confidence_minus_random_symmetric_seed_405Reported0.298066[0.285532, 0.310298]0.0059972,000
H1Self-confidence vs random review · Symmetric · Seed 406h1_self_confidence_minus_random_symmetric_seed_406Reported0.301792[0.288646, 0.314714]0.0059972,000
H3Symmetric vs confusion-targeted corruption · Seed 404h3_symmetric_minus_confusion_seed_404Reported0.008595[0.000099, 0.016694]0.0479762,000
H3Symmetric vs confusion-targeted corruption · Seed 405h3_symmetric_minus_confusion_seed_405Reported0.007071[-0.000892, 0.014935]0.0479762,000
H3Symmetric vs confusion-targeted corruption · Seed 406h3_symmetric_minus_confusion_seed_406Reported0.011085[0.002823, 0.019583]0.0104952,000
H3Symmetric vs instance-dependent corruption · Seed 404h3_symmetric_minus_instance_dependent_seed_404Reported0.256775[0.242522, 0.270189]0.0029992,000
H3Symmetric vs instance-dependent corruption · Seed 405h3_symmetric_minus_instance_dependent_seed_405Reported0.261788[0.248217, 0.275018]0.0029992,000
H3Symmetric vs instance-dependent corruption · Seed 406h3_symmetric_minus_instance_dependent_seed_406Reported0.265437[0.251198, 0.279251]0.0029992,000
H5Fixed hybrid vs self-confidence · Confusion-targeted · Seed 404h5_hybrid_minus_self_confidence_confusion_seed_404Reported0.052548[0.046700, 0.058534]0.0059972,000
H5Fixed hybrid vs self-confidence · Confusion-targeted · Seed 405h5_hybrid_minus_self_confidence_confusion_seed_405Reported0.054609[0.048486, 0.061070]0.0059972,000
H5Fixed hybrid vs self-confidence · Confusion-targeted · Seed 406h5_hybrid_minus_self_confidence_confusion_seed_406Reported0.054104[0.048107, 0.060500]0.0059972,000
H5Fixed hybrid vs self-confidence · Group-conditional · Seed 404h5_hybrid_minus_self_confidence_group_conditional_seed_404Reported0.060268[0.053809, 0.066753]0.0059972,000
H5Fixed hybrid vs self-confidence · Group-conditional · Seed 405h5_hybrid_minus_self_confidence_group_conditional_seed_405Reported0.065135[0.058752, 0.071457]0.0059972,000
H5Fixed hybrid vs self-confidence · Group-conditional · Seed 406h5_hybrid_minus_self_confidence_group_conditional_seed_406Reported0.065502[0.058992, 0.071871]0.0059972,000
H5Fixed hybrid vs self-confidence · Instance-dependent · Seed 404h5_hybrid_minus_self_confidence_instance_dependent_seed_404Reported0.020983[0.018483, 0.023472]0.0059972,000
H5Fixed hybrid vs self-confidence · Instance-dependent · Seed 405h5_hybrid_minus_self_confidence_instance_dependent_seed_405Reported0.020983[0.018483, 0.023472]0.0059972,000
H5Fixed hybrid vs self-confidence · Instance-dependent · Seed 406h5_hybrid_minus_self_confidence_instance_dependent_seed_406Reported0.020983[0.018483, 0.023472]0.0059972,000
H5Fixed hybrid vs self-confidence · Symmetric · Seed 404h5_hybrid_minus_self_confidence_symmetric_seed_404Reported0.062592[0.056457, 0.068956]0.0059972,000
H5Fixed hybrid vs self-confidence · Symmetric · Seed 405h5_hybrid_minus_self_confidence_symmetric_seed_405Reported0.064366[0.057757, 0.070856]0.0059972,000
H5Fixed hybrid vs self-confidence · Symmetric · Seed 406h5_hybrid_minus_self_confidence_symmetric_seed_406Reported0.064656[0.058361, 0.071218]0.0059972,000
H6Pathology encoder vs ImageNet · Seed 404h6_pathology_minus_imagenet_seed_404UnavailableUnavailableUnavailableUnavailable0
H6Pathology encoder vs ImageNet · Seed 405h6_pathology_minus_imagenet_seed_405UnavailableUnavailableUnavailableUnavailable0
H6Pathology encoder vs ImageNet · Seed 406h6_pathology_minus_imagenet_seed_406UnavailableUnavailableUnavailableUnavailable0
H7Target-highlighted vs context · Seed 404h7_highlighted_minus_context_seed_404Reported0.002997[-0.005062, 0.010787]0.7226392,000
H7Target-highlighted vs context · Seed 405h7_highlighted_minus_context_seed_405Reported-0.000004[-0.008175, 0.008265]0.9645182,000
H7Target-highlighted vs context · Seed 406h7_highlighted_minus_context_seed_406Reported-0.005371[-0.013518, 0.003130]0.9645182,000

Statistical note. Holm-adjusted p-values are the registered one-sided tests. The saved 95% confidence intervals are percentile-bootstrap summaries based on whole-group resampling.

Research sources

Grouping follows the strongest identifiers available per dataset. No result is promoted beyond the grouping and reference-label evidence that the source actually provides.

Reproducibility identities

Accepted run
20260727T133947.089370Z_pannuke_primary_orphan_recovery
Artifact root SHA-256
8c1c7b277d96889dc4fb45aee282e77e3d351f687990e03e6b57ec5f2313c7e4
Stage attestation SHA-256
5af827544502fbdf688a73916ec58b5dac0984c5a682a33ce6dfc97538228871
QC overlay SHA-256
a1bd87dd397417d711d1d4937429eae5f5d972d3fa6ffa27a45129339587f10a

The same system transferred under controlled noise, but natural-case proof is still missing.

Later evidence changes the overall assessment without rewriting the adverse PanNuke H4 result. Each study answers a different question and keeps its own frozen boundary.

NuCLS is the closest completed test of genuine multi-rater disagreement. Its five-patient result did not satisfy the frozen ranking gate and the guided intervention made downstream macro-F1 worse. MoNuSAC then showed strong controlled retrieval but no statistically supported downstream gain or complete class safety.

The selected AANCA candidate was next recorded internally as frozen before PUMA outcomes. On a previously unused histopathology source it passed all seven internally pre-specified retrieval, downstream, direction, convergence and class-safety gates under controlled corruption. The 144/62 development/final partition was created by AANCA from the 206 public PUMA ROIs; it was not the official hidden PUMA challenge test. The downstream flag_exclude arm omitted the highest-ranked 5% of training rows without expert review or relabelling. This is meaningful controlled-noise transfer evidence, but PUMA does not provide paired natural pre/post expert decisions and therefore cannot prove pathologist-error detection.

Public-history limit. The PUMA protocol, configuration and result first appeared together in public commit c5bd44193b2abd67bc7e7f1bd9384aa87435d500, so public Git history does not independently verify the intended pre-outcome timing. The PUMA verifier is a project-coupled evidence-readback script that recomputes metrics from saved predictions but does not retrain all 44 models. It is not third-party validation or a second image-to-result replication.

NuCLS natural multi-rater disagreement

On 5 patient groups and 811 nuclei, average precision was 0.073489. Precision at the 5% review budget was 0.097561, but its difference-from-prevalence interval crossed zero. Guided correction reduced macro-F1 by 0.014633, with a 95% interval [-0.026683, -0.002415]. The frozen natural-data claims were not supported.

MoNuSAC controlled external benchmark

The frozen queue found 1,035 of 2,961 injected changes: precision 0.698852 versus 0.556009 for matched random review, with a positive difference interval [0.099181, 0.188491]. Downstream macro-F1 changed by +0.005526 [-0.001506, +0.012833], so the benefit and class-safety gates failed. The prescribed action remained retain_uncorrected.

PUMA internally frozen new-source controlled confirmation

The candidate was recorded internally as frozen before PUMA outcomes; public Git history does not independently verify that timing. It was evaluated on 62 held-out case/ROI groups from an AANCA-defined 144/62 split of the 206 public ROIs, not the official hidden PUMA challenge test set. Queue precision was 0.537739 versus 0.214379; the difference was +0.323359 [+0.259251, +0.384944]. Candidate macro-F1 was 0.646310, improving on unchanged labels by +0.006426 [+0.003657, +0.009365] and matched random by +0.008067 [+0.004093, +0.011947]. All seven internally pre-specified gates passed. The flag_exclude arm omitted the highest-ranked 5% of controlled training rows; they were not reviewed, corrected or automatically relabelled by an expert. This supports controlled-noise transfer only.

Robustness and the remaining safety boundary

All 9 of 9 post-confirmation PUMA stress scenarios had a positive aggregate downstream lower bound, but only 1 passed every class safeguard. The puma_audit_time_label_sensitivity_v1 observed-label fold-allocation sensitivity passed all seven gates, but it was run after PUMA outcomes were open. Natural paired pre/post NuCLS evidence was unavailable, so the binding action is still retain_uncorrected.

The source masks remained untouched.

The validator measured cross-class overlaps and unlabeled regions, retained them in provenance and applied one frozen eligibility policy. It never arbitrated an overlap class or reconstructed the supplied background.

Before any primary split is frozen, the inspector verifies folds, shapes, channels, class order, instance identifiers and representative overlays. Cross-class overlaps, void pixels and affected instances are recorded rather than silently corrected.

The overlay is evidence about ingestion and representation, not improved segmentation. Eligibility flags may follow the frozen policy, but the raw masks and source annotations remain unchanged.

7,901validated PanNuke patches
4,318cross-class overlap pixels
10,486,091unlabeled / void pixels
1,411overlap-touching instances flagged
Deterministic quality-control overlay1512 x 3840 px · source preview
Cropped representative preview · normal, overlap, void and exclusion casesOpen the original image ↗
Deterministic PanNuke quality-control overlays showing images and mask boundaries

The design limits outcome-informed model selection.

The model, split, feature and statistical rules were frozen before outcome interpretation. Because outcomes were later exposed during recovery, this accepted result is reported as exploratory rather than as an untouched confirmatory analysis.

Source groups stay together, each audit score is out of fold, the final reference fold is withheld from selection, and every label state remains separate. The matrix retained all 36 comparison entries; unavailable optional pathology cells were not replaced with estimates.

Outcome exposure during technical recovery does not change the stored measurements, but it narrows what can responsibly be claimed. The accepted primary analysis is therefore permanently described as amended and exploratory.

185/185required primary cells completed
0failed required cells
36preregistered comparison entries retained
2,000paired group-bootstrap iterations
Group-safe splitting

Every split uses group_id, at least at source-patch level, never individual nuclei.

Out-of-fold ranking

Primary model-based audit scores come from predictions made for source groups excluded from model fitting.

Untouched final reference

The final reference fold is uncorrupted and unavailable for model selection, calibration or review-budget tuning.

Separate label states

pre_corruption_label, observed_label, is_injected_corruption and corruption metadata remain distinct.

What this benchmark supports, and what it does not.

A technically strong audit can still be limited in scope. The boundary below is part of the result, not a disclaimer added afterward.

The controlled positive event is an injected label change. Performance against it measures retrieval of that injected process, not naturally occurring annotation inconsistency, expert disagreement or biological truth. The positive PUMA result is a new-source controlled transfer result, while the later stress analysis and observed-label sensitivity are exploratory because PUMA outcomes were already open.

Grouping strength differs by source: the primary PanNuke study guarantees source-patch separation, NuCLS uses patient groups, MoNuSAC uses patient groups and PUMA uses one ROI per case. These safeguards reduce leakage but do not transform a final expert reference label into guaranteed biological truth.

Natural and operational validity still require newly recruited blinded reviewers, ambiguity and abstention labels, a frozen policy evaluated on untouched patients or whole slides, and a prospective multi-site comparison of work with and without AANCA.

Supported by current evidence

  • Group-safe rankings retrieve injected class-label changes more efficiently than matched random review.
  • The frozen current candidate transferred to new PUMA images and improved controlled-noise downstream macro-F1 with positive whole-group intervals.
  • Every displayed conclusion is read from checksum-bound machine-readable evidence.

Not established

  • That a naturally occurring annotation is wrong, that a pathologist made an error or that model disagreement is biological truth.
  • That the intervention is uniformly class-safe across realistic corruption patterns; only one of nine stress scenarios passed every class safeguard.
  • Prospective workflow benefit, multi-site generalisation, clinical utility or permission to alter natural source labels automatically.

AANCA is externally evaluated, but not yet confirmatory or ready for real-use claims.

Completion labels describe which governed evaluations ran. They do not turn a mixed result into efficacy.

The current project has reached PRIMARY_STUDY_COMPLETE, EXTERNAL_VALIDATION_COMPLETE and DEMO_COMPLETE. It has not reached CONFIRMATORY_COMPLETE. The binding natural-data action is retain_uncorrected: AANCA may prioritise cases for qualified review, but it may not silently exclude, relabel or overwrite them.

The next research phase is provisionally named AANCA v2. This is a prospective evidence programme for the same core auditing system, not a retroactive rename of the existing results. Opened PanNuke, NuCLS, MoNuSAC and PUMA outcomes may inform diagnosis of weaknesses, but they cannot serve as the untouched final confirmation for the next claim.

1 · Natural reference

Recruit independent blinded pathologists on new cases and preserve agreement, disagreement, ambiguity, abstention and insufficient-context outcomes instead of forcing one truth label.

2 · Measured utility

Develop the review queue inside nested patient- or WSI-group cross-fitting using both inconsistency probability and conservatively estimated downstream benefit.

3 · Prospective freeze

Freeze one representation, queue, intervention, review budget, class-safety rule and analysis plan before inspecting any new confirmation outcome.

4 · Untouched confirmation

Run one-shot external validation on new patient or WSI groups. Ranking, downstream confidence intervals, every-class safety and convergence must all pass together.

5 · Real workflow

Compare blinded multi-site review with and without AANCA, measuring time, agreement, accepted corrections, downstream performance and failure modes.

Promotion boundary

Only the full sequence can support a claim about realistic natural-case improvement. Until then, AANCA remains a non-diagnostic expert-review prioritisation prototype.

Read the evidence first; run the software when a deeper check is needed.

The checked-in article opens without a dataset, model run or GPU. present_demo.py --verify-only verifies the thirteen-file presentation package and its current PanNuke, NuCLS, MoNuSAC and PUMA summaries; it does not retrain a model or recalculate a scientific result. The separate synthetic path exercises the portable software workflow.

The public repository retains frozen protocols, configs, compact results and arrays for scoped evidence verification. The primary, NuCLS and MoNuSAC scripts independently recalculate their stated saved evidence. The PUMA readback imports maintained helpers and checks stored predictions rather than independently retraining the 44 models. Full image-to-result replication still requires lawfully obtained source datasets and appropriate compute; the project never downloads protected data silently or relaxes scientific gates when an input is unavailable.

github.com / JaqwilkAANCA
View repository ↗

The repository contains the scientific specification, maintained source code, frozen study configs, scoped verification scripts, compact evidence and this checksum-verifiable presentation. Licensed raw datasets, reusable local embeddings and unretained historical checkpoints are deliberately excluded.

Python 3.12group-safe OOFimmutable evidenceresearch only

Verify and open locally

Check the package, serve it locally and open the article.

git clone https://github.com/Jaqwilk/AANCA.git
cd AANCA
python scripts/present_demo.py

Check without a browser

Validate the presentation package without recomputing the study.

python scripts/present_demo.py --verify-only

Run the synthetic smoke path

Exercise the software with deterministic synthetic data.

uv sync --dev
uv run histo-audit doctor
uv run histo-audit data generate-synthetic --config configs/smoke.yaml
uv run histo-audit experiment smoke --runs-root artifacts/smoke_runs

Real PanNuke runs require a lawful local dataset and the governed setup in DATASET_SETUP.md.

Research and implementation by Natan Smogór.

Natan Smogór is the author and developer of AANCA, a non-diagnostic research prototype for prioritising potentially inconsistent nucleus annotations for expert review.

The project combines a frozen scientific specification, controlled data interventions, group-safe evaluation, machine-readable provenance and an inspectable presentation layer. Its purpose is to demonstrate a reproducible research workflow while keeping the final interpretation with qualified experts.

AI tools assisted throughout the project with planning, code drafting and iterative implementation. Natan Smogór directed the work, reviewed and revised the AI-assisted outputs, made the final scientific and engineering decisions, and remains responsible for the code, analysis and claims.