Skip to content
Yinheng ZHU

Decoupled Single-Mask Annotation Noise Detection via Cross-Sectional Patch Self-Consistency

Yinheng Zhu1 and Xiaowei Xu2,✉

1 Tsinghua University, Beijing, China
2 Guangdong Provincial People's Hospital, Guangzhou, China
✉ corresponding author

Medical Image Computing and Computer Assisted Intervention (MICCAI) 2026
Oral Poster


Annotation noise is not random

Vascular CT is usually annotated once per scan, because repeated expert labelling is costly. A single mask leaves nothing to compare it against, so it is hard to say which parts of it are wrong. We proposed a way to find them from that one mask, with no second rater and without touching network training. It is cheap enough to run over a whole dataset: across about three million vessel cross-sections, the errors are not random.

We found that annotation noise correlates strongly with the direction a vessel runs. Below is the noise rate over vessel direction: how often a cross-section is flagged, as a function of which way it points. The six poles are the imaging axes. The rate dips there, and those are the cleanest annotations. The bulges between them are vessels running diagonally through the scanner.

Direction sphere Drag to rotate · hover the surface

Loading direction data…

The noise rate over vessel tangent direction, binned 80×40 and encoded by both radius and colour. A tangent and its opposite denote the same direction and are binned together. Cell solid angle shrinks as sin φ, so counts are pooled over a neighbourhood that widens toward the poles and shrunk toward the global rate with a Beta prior; without this the sparse polar cells are too noisy to display. Hovering reports the observed rate alongside the smoothed one. Sample density shows that the oblique regions are not sparsely sampled.

Measured as the angular distance θ to the nearest imaging axis, the rate rises close to monotonically across the range, from 0.44% at θ ≈ 0° to 2.24% at θ ≈ 55°. Spearman ρ = −0.20; Cochran–Armitage trend test z = −33.31; both p < 0.001.

Rate by angle Drag across the plot to filter the walker · click to clear
Noise rate per bin with Wilson 95% intervals. Brushing a range restricts which cross-sections the walker below will jump to.

And not only orientation

Annotation noise also correlates with cross-sectional area and with intensity, and these two share one cause. Cross-sections under 2 mm² show markedly elevated noise rates, declining rapidly beyond 5 mm². The rate also peaks below 0 HU, at approximately 5–6%, and drops through the typical lumen range of 50–200 HU. Thin vessels occupy only a few pixels in axial slices, and low contrast leaves less separation between vessel wall and background. Both trends reflect the same condition: reduced boundary discriminability at annotation time.

Area and contrast Drag across a plot to filter the walker · click to clear
The wide confidence bands at the extremes reflect sparse samples: the smallest area bin holds few cross-sections.

Our explanation for these findings is that it may correlate with the annotation tools. Annotation is done slice by slice, in tools such as ITK-SNAP. A vessel running along an imaging axis appears in that slice as a circular cross-section, whereas an oblique vessel appears as an elongated ellipse with a diffuse boundary. On this reading the pattern is not anatomical but an artefact of the annotation interface. Hemodynamic workflows corroborate this by annotating directly on centreline cross-sections.

The principle and auditable evidence

A single pair of cross-sections, drawn from two different scans, states the problem the method addresses.

Two coronary trees from different scans; a cross-section is taken from each; the two cross-sections are the same while the two expert masks differ substantially Same image. Different mask.
Two scans from ImageCAS, each with one cross-section taken where the dotted line points. The two cross-sections fall within the retrieval tolerance of 30 dB PSNR, at which a pair is visually indistinguishable. The masks drawn on them differ substantially, so at least one of the two annotations is unreliable. Neither scan contains anything that would reveal this, since each was annotated once by one rater.

Existing approaches fall into two families. Multi-rater fusion rests on one principle: the same image should receive the same mask, and where it does not, there is annotation noise. Training-coupled robust learning instead infers noise from optimisation dynamics, and so cannot separate a mislabelled region from a merely hard one. The first principle requires either (a) that the same image has been annotated more than once, which is the cost we cannot afford, or (b) that identical images already exist within the dataset.

Tubular anatomy provides option (b). Sampled perpendicular to the centreline, cross-sections recur along the same vessel and across different subjects, so the dataset already contains its own near-repeats without any additional annotation. The principle then needs one word relaxed, from the same image to a near-identical one:

dI(Ii, Ij) ≤ εI  ⇒  dM(Mi, Mj) ≤ εM

Image distance dI is mean squared error; mask distance dM is one minus intersection over union. Each patch is compared with every neighbour retrieved within εI, and the median of those pair scores gives its patch-level score Ri: a signed z-score of mask disagreement, calibrated against how much disagreement is normal for pairs of that similarity. A patch whose mask disagrees with most of its retrieved neighbours is driven to Ri << 0.

The criterion is computed directly from the image–mask pair rather than from network behaviour, so each flagged region is surfaced together with the retrieved neighbours that produced its score. This is what the work set out to provide: in a single-rater dataset, annotation noise identified decoupled from training, on auditable evidence. The demo below steps along one centreline and shows the retrieved neighbours at each cross-section.

Centreline walker Drag the trace, or use ← → · orbit the mesh

Loading scan…

current cross-section

Retrieved neighbours that set this score, ordered by mask disagreement. A trailing * marks one from a different scan.

← mask agreesmask disagrees →

Every point is one cross-section, in centreline order; gaps are branch breaks. The trace is Ri: lower is worse, and the dashed line marks the −3 threshold. Neighbour counts are small because few patches qualify under the retrieval tolerance, and none that fail it are included.

The score is graded rather than binary: it separates severe disagreement from mild boundary shifts, and it also identifies clean annotation as clean.

Score spectrum Each column is a slice of the score · each row one pair

Loading spectrum…

Each pair is a scored cross-section beside the retrieved neighbour that disagrees with it most. At the left the images are near-identical while the masks differ drastically; toward the right the discrepancies become local boundary shifts; at the right end the masks agree. Examples are drawn from the whole dataset rather than from a single scan, since the most extreme bin is sparsely populated and spread across many subjects.

How it works

The pipeline has three stages, none of which involves training a network.

Pipeline: cross-sectional patches sampled along centreline frames, neighbours retrieved across subjects, scores aggregated into a scan-level quality map
  • Sample patches along the centreline.
  • 103 volumes produce ∼3×106 patches.
  • Bishop frames keep them stable.
  • Operationalise the principle as a statistical test.
  • Implemented with a vector search engine to accommodate 106+6 patch pairs.
  • Spread patch residuals onto voxels to obtain a quality map for loss weighting.
Left: patches are extracted along centreline frames and visually similar ones retrieved across subjects. Middle: each mask is compared with those of its intensity-equivalent neighbours. Right: pair- and patch-level scores aggregate into a scan-level quality map.

Three choices deserve comment. The frames are torsion-free: a frame that twists as it travels would rotate the patch with it, and a rotated cross-section is no longer retrievable as near-identical, which would cost the method the pairs it depends on. The retrieval tolerance εI = 10−3 MSE corresponds to about 30 dB PSNR, the level at which two patches stop being distinguishable by eye, which is what licenses treating a retrieved neighbour as the same image. And mask disagreement is not compared against a fixed threshold: pairs are binned by image distance and calibrated within each bin, so a pair is judged against the disagreement that is normal at its own level of similarity.

Failure cases and limitations

The dominant failure pattern is false negatives. Noise goes undetected when no sufficiently similar counterpart is retrieved: patches that are near-identical only up to a rotation, or up to a window-level shift, are missed by raw-MSE retrieval. Centreline and frame errors have the same effect. All of these reduce recall rather than precision, which is the trade the design prioritises. Recall is recoverable by denser sampling of rotations and shifts, at higher computational cost.

The method detects rather than corrects noise, assumes tubular anatomy, and has been evaluated on ImageCAS alone.

Two patches near-identical up to a rotation, whose masks disagree Two patches near-identical up to a window-level shift, whose masks disagree
Undetected noise. Left: near-identical up to a rotation. Right: near-identical up to a window-level shift. Both pairs are real annotation disagreements that the retrieval does not surface.

References

How to cite

@inproceedings{zhu2026hiddennoise,
  title     = {Decoupled Single-Mask Annotation Noise Detection via
               Cross-Sectional Patch Self-Consistency},
  author    = {Zhu, Yinheng and Xu, Xiaowei},
  booktitle = {Medical Image Computing and Computer Assisted Intervention
               (MICCAI)},
  year      = {2026},
  eprint    = {2607.05965},
  archivePrefix = {arXiv},
}

Acknowledgements

Experiments use the ImageCAS coronary CT angiography dataset.