MedVision Logo MedVision: Benchmarking Quantitative Medical Image Analysis

1School of Informatics, University of Edinburgh
2School of Engineering, University of Edinburgh
3Queen Mary University of London
All posts

MedVision v1.4.0: 50× More Tumor/Lesion Measurements and Their Annotation Recall

Summary. Dataset v1.4.0 rebuilds the tumor/lesion (T/L) size annotations of all 12 tumor and lesion collections. Clusters are selected by a physical size floor in millimetres instead of a raw pixel count, a containment test that discarded rotated ellipses is removed, and four guards bracket the ellipse fit. Published measurements increase from 75,840 to 3,801,540 (50×). This release also quantifies annotation recall for the first time: of 6,951,667 outlined clusters, 3,801,540 (0.547) carry a measurement, and recall reaches 0.992 for clusters at least 20 mm long.

Background

Tumor and lesion size is a quantitative clinical endpoint. It informs staging, treatment selection, and response assessment across successive scans. MedVision evaluates whether vision language models can recover that quantity from medical images, which requires a reference standard: a large image corpus in which the physical size of each tumor or lesion is known.

Measurement procedure

Every volume in these collections carries expert segmentation masks of diseased tissue. On each slice, every connected component of a target-label mask (a cluster) is fitted with an ellipse, and the major and minor axis lengths are recorded in millimetres, mirroring how a radiologist sizes a lesion. Those lengths are the reference values against which a model is graded.

Because the fit is automatic, the rules that decide which clusters are measured and which fits are trustworthy determine the composition of the dataset. v1.4.0 revises those rules.

Changes in v1.4.0

No measurement published by an earlier version was incorrect. The earlier rules were restrictive in ways that left many measurable lesions out, and the extent of that loss had never been quantified.

1. Selection in millimetres rather than pixels. A pixel is not a fixed physical size: it varies between scanners and between viewing directions within one volume. On a case of 0.977/0.977/3.0 mm spacing, the earlier 20 px floor cut at approximately 4.9 mm in the axial plane but 8.6 mm in the sagittal plane, a 3.07× difference in area behind a single number. v1.4.0 measures a cluster when its fitted major axis clears max(2.0 mm, 2× the coarser in-plane spacing), a resolution floor applied identically to every image.

2. Removal of the orientation-sensitive containment test. The earlier rule required the fitted ellipse to lie inside a box around the cluster, which tests shape and rotation rather than size: a tilted ellipse protrudes from an axis-aligned box regardless of fit quality. Most of the 50× growth is attributable to this removal rather than to the lower floor.

3. Four guards on the fit. Fitting millions of contours automatically admits rare degenerate solutions. v1.4.0 rejects contours under 5 points, non-finite conics, minor axes thinner than one voxel, and major axes exceeding 1.5× the cluster’s own bounding-box diagonal. The logged guards rejected 95,800 fits corpus-wide.

Measurement distribution

Box plot of major axis distributions for all 38 tumor and lesion types in MedVision v1.4.0

Distribution of the major axis for all 3,801,540 measurements. Each row is one of the 38 tumor or lesion labels; the box spans the interquartile range and the whiskers the central 90%, in millimetres.

The median major axis spans 7.7 mm (non-enhancing brain tumor core, BraTS24-GLI) to 43.3 mm (kidney tumor, KiTS23), so a single expected size does not generalise across diseases. The largest measurement in the corpus, 540.1 mm.

Annotation recall

Recall is measured clusters divided by all clusters, with the denominator obtained by recounting the expert masks directly rather than by reading the annotation plan. Across the 12 collections the masks contain 6,951,667 clusters, of which 3,801,540 (0.547) carry a measurement.

Recall by cluster size in pixels

Per bin Cumulative recall (≤ t) Recall above threshold (> t)
Cluster sizeClustersMeasuredRecallCluster sizeClustersMeasuredRecallCluster sizeClustersMeasuredRecall
1–2 px1,872,14900.000≤2 px1,872,14900.000>2 px5,079,5183,801,5400.748
3–5 px831,75134,3470.041≤5 px2,703,90034,3470.013>5 px4,247,7673,767,1930.887
6–10 px665,002324,7160.488≤10 px3,368,902359,0630.107>10 px3,582,7653,442,4770.961
11–20 px649,587551,3370.849≤20 px4,018,489910,4000.227>20 px2,933,1782,891,1400.986
21–50 px756,110723,7790.957≤50 px4,774,5991,634,1790.342>50 px2,177,0682,167,3610.996
51–100 px490,959485,4180.989≤100 px5,265,5582,119,5970.403>100 px1,686,1091,681,9430.998
101–500 px1,003,733999,9610.996≤500 px6,269,2913,119,5580.498>500 px682,376681,9820.999
501–1000 px346,583346,2420.999≤1000 px6,615,8743,465,8000.524>1000 px335,793335,7401.000
>1000 px335,793335,7401.000All6,951,6673,801,5400.547

The first block is the recall of each bin in isolation. The second accumulates the bins upward from the smallest, the cumulative recall in the usual sense, and closes on the corpus total. The third pools every cluster above a threshold, which is the quantity a user of the dataset sees when selecting lesions above a size floor; the second and third blocks at one threshold partition the corpus exactly. Read from the third block, recall is 0.961 for clusters >10 px and 0.986 for clusters >20 px: the floor and the fit guards remove sub-resolution fragments, not measurable lesions. The fragment shares come from the second block, converted from cluster counts to unmeasured counts by subtracting the measured column from the clusters column in the same row:

The unmeasured remainder is dominated by fragments: 95.5% of unmeasured clusters are ≤10 px and 84.7% are ≤5 px. The corpus value of 0.547 therefore counts components, not clinically measurable lesions. The pixel table below supports those shares, and the two physical stratifications that follow separate the two quantities; each table partitions the same 6,951,667 clusters and reproduces the same margins.

unmeasured, all sizes = 6,951,667 - 3,801,540 = 3,150,127
unmeasured ≤10 px     = 3,368,902 -   359,063 = 3,009,839    ->  3,009,839 / 3,150,127 = 0.955
unmeasured ≤5 px      = 2,703,900 -    34,347 = 2,669,553    ->  2,669,553 / 3,150,127 = 0.847

Clusters of 1–2 px, over a quarter of the corpus, cannot yield the 5-point contour the ellipse fit requires, so their recall is 0 by construction rather than by threshold.

Recall by major axis length

Length is the maximum Feret diameter of a cluster in millimetres, the mask-side proxy for the fitted major axis on which the selection rule acts.

Per bin Cumulative recall (≤ t) Recall above threshold (> t)
Cluster lengthClustersMeasuredRecallCluster lengthClustersMeasuredRecallCluster lengthClustersMeasuredRecall
≤2 mm2,232,4271,9520.001≤2 mm2,232,4271,9520.001>2 mm4,719,2403,799,5880.805
2–5 mm974,881409,4120.420≤5 mm3,207,308411,3640.128>5 mm3,744,3593,390,1760.905
5–10 mm955,724694,6830.727≤10 mm4,163,0321,106,0470.266>10 mm2,788,6352,695,4930.967
10–20 mm1,055,235976,1050.925≤20 mm5,218,2672,082,1520.399>20 mm1,733,4001,719,3880.992
20–50 mm1,263,2411,250,2680.990≤50 mm6,481,5083,332,4200.514>50 mm470,159469,1200.998
50–100 mm413,816412,8490.998≤100 mm6,895,3243,745,2690.543>100 mm56,34356,2710.999
>100 mm56,34356,2710.999All6,951,6673,801,5400.547

Read from the third block, recall is 0.967 for clusters >10 mm and 0.992 for clusters >20 mm.

The same table sizes the unmeasured remainder in physical units, and the result is that very few unmeasured clusters are large lesions: of the 3,150,127 unmeasured clusters, 88.8% are ≤5 mm, 97.0% are ≤10 mm, and only 14,012 (0.44%) are longer than 20 mm. Each share is derived as before — clusters minus measured, in the row at that threshold, over the corpus-wide unmeasured total:

unmeasured, all lengths = 6,951,667 - 3,801,540 = 3,150,127
unmeasured ≤5 mm        = 3,207,308 -   411,364 = 2,795,944  ->  2,795,944 / 3,150,127 = 0.888
unmeasured ≤10 mm       = 4,163,032 - 1,106,047 = 3,056,985  ->  3,056,985 / 3,150,127 = 0.970
unmeasured >20 mm       = 1,733,400 - 1,719,388 =    14,012  ->     14,012 / 3,150,127 = 0.0044

Recall by cluster area

Area is the physical area of the cluster on its slice in mm², derived from the same recount.

Per bin Cumulative recall (≤ t) Recall above threshold (> t)
Cluster areaClustersMeasuredRecallCluster areaClustersMeasuredRecallCluster areaClustersMeasuredRecall
≤2 mm²1,684,011640.000≤2 mm²1,684,011640.000>2 mm²5,267,6563,801,4760.722
2–5 mm²696,04140,0930.058≤5 mm²2,380,05240,1570.017>5 mm²4,571,6153,761,3830.823
5–10 mm²526,878237,7020.451≤10 mm²2,906,930277,8590.096>10 mm²4,044,7373,523,6810.871
10–20 mm²545,631371,5690.681≤20 mm²3,452,561649,4280.188>20 mm²3,499,1063,152,1120.901
20–50 mm²725,574530,0740.731≤50 mm²4,178,1351,179,5020.282>50 mm²2,773,5322,622,0380.945
50–100 mm²582,689467,4690.802≤100 mm²4,760,8241,646,9710.346>100 mm²2,190,8432,154,5690.983
100–500 mm²1,347,0051,311,2160.973≤500 mm²6,107,8292,958,1870.484>500 mm²843,838843,3530.999
500–1000 mm²428,510428,1200.999≤1000 mm²6,536,3393,386,3070.518>1000 mm²415,328415,2331.000
>1000 mm²415,328415,2331.000All6,951,6673,801,5400.547

Read from the third block as before, recall reaches 0.983 for clusters >100 mm² and 0.999 for clusters >500 mm². Area does not determine the major axis: at a fixed area an elongated cluster has a longer major axis than a compact one and is more likely to clear the floor, while the thinnest clusters are removed by the sub-voxel minor-axis guard. Length, the variable the rule acts on, is therefore the sharper predictor of recall, and the area view describes the same corpus with shapes mixed together.

The unmeasured remainder concentrates at the bottom of this ladder as it does on the other two: 83.5% of the 3,150,127 unmeasured clusters are ≤10 mm² and 89.0% are ≤20 mm², derived as before:

unmeasured, all areas = 6,951,667 - 3,801,540 = 3,150,127
unmeasured ≤10 mm²    = 2,906,930 -   277,859 = 2,629,071  ->  2,629,071 / 3,150,127 = 0.835
unmeasured ≤20 mm²    = 3,452,561 -   649,428 = 2,803,133  ->  2,803,133 / 3,150,127 = 0.890
Recall of tumor and lesion measurements above each cluster-area and cluster-length threshold, one curve per collection, MedVision v1.4.0

Recall above threshold, against cluster area (left) and cluster length (right), one curve per collection. Each point pools every cluster larger than the threshold, the quantity read from the third block of the tables. Both panels stratify the same 6,951,667 clusters. Recall rises with the floor in every collection and is at least 0.972 in every collection for clusters >20 mm long.

Implications for users

The images are unchanged; only the T/L measurements were rebuilt. They are larger in number, selected consistently across scans and viewing directions, and free of degenerate values. Every earlier annotation version (<1.4.0) remains published unchanged, so previous results stay reproducible.

One caveat applies to comparisons. Case counts per split are unchanged, but train/test membership changes for six collections (autoPET-III, BraTS24, HNTSMRG24, KiPA22, KiTS23, MSD), whose earlier splits had been force-aligned to v1.0.0. A v1.4.0 test metric should not be compared against a pre-1.4.0 one on those six.

The full release note documents every rule, number, and verification behind this post.

Acknowledgements

This work was supported by UK Research and Innovation (grant EP/S02431X/1) through the UKRI Centre for Doctoral Training in Biomedical AI at the School of Informatics, University of Edinburgh. We also acknowledge the support of the Edinburgh International Data Facility (EIDF) and the Data-Driven Innovation Programme at the University of Edinburgh.

BibTeX

@misc{yao2026medvisionbenchmarkingquantitativemedical,
    title={MedVision: Benchmarking Quantitative Medical Image Analysis}, 
    author={Yongcheng Yao and Yongshuo Zong and Raman Dutt and Yongxin Yang and Sotirios A Tsaftaris and Timothy Hospedales},
    year={2026},
    eprint={2511.18676},
    archivePrefix={arXiv},
    primaryClass={cs.CV},
    url={https://arxiv.org/abs/2511.18676}, 
}