매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

A Strong Linear Baseline for Whole-Heart Cardiac Shape Completion on CT, with an Open Eleven-Structure Statistical Shape Model

arXiv:2608.199322026-08-21

심장 CT 11개 구조물을 하나로 잇는 통계적 형상 모델, 복잡한 딥러닝보다 간단한 수식이 더 정확했다

심장을 찍은 CT 데이터마다 심방이개, 폐정맥, 대정맥 같은 구조물이 있기도 하고 없기도 해서 서로 다른 병원 데이터를 합치기 어렵다는 문제에서 출발했다. 연구팀은 383건의 CT에서 11개 심장 구조물을 동일한 정점 대응 관계로 통일한 형상 모델을 만들고, 일부 구조물만 보고 나머지를 복원하는 여러 방법을 같은 조건에서 비교했다. 그 결과 닫힌 형태의 수식으로 계산하는 조건부 가우시안 추정기가 그래프 기반 딥러닝 모델보다 오차가 더 작았다.

무엇을 했나

  1. 공개된 심장 형상 데이터셋 중 심방이개, 폐정맥, 대정맥 끝부분을 별도의 온전한 표면 덩어리로 포함한 것이 없어, 연구팀이 383건의 CT로 11개 구조물을 11571개 정점으로 서로 대응시킨 통계적 형상 모델(SSM)을 새로 만들어 공개했다
  2. 일부 구조물만 관찰된 상태에서 나머지를 예측하는 완성(completion) 문제를 두고, 수식으로 정확히 풀리는 조건부 가우시안 추정기와, 마스크 조건을 학습한 그래프 변분 오토인코더(그래프 β-VAE), 최근접 이웃 검색 등을 동일한 76건의 내부 평가셋과 동일한 조건으로 비교했다
  3. 조건부 가우시안 추정기는 평균 정점 오차 3.717mm를 기록해, 그래프 β-VAE의 5.248mm, 최근접 이웃 검색의 8.931mm보다 낮았고, 그래프 β-VAE와의 차이는 1.531mm(95% 신뢰구간 1.384~1.711mm)였다
  4. 58건의 외부 CT(CARE2026)와 20건의 다른 공개 벤치마크(MM-WHS)에서 전문가가 직접 라벨링한 일부 구조물로 검증했을 때도 같은 순서(조건부 가우시안이 더 정확함)가 유지됐지만, 대동맥 등 4개 구조물은 참조 데이터가 충분히 정확하지 않아 비교 대상에서 제외됐다
  5. 연구팀은 이 모델과 완성 방법이 정렬된 CT 데이터를 하나로 통합하는 연구용 도구이며, 임상 진단이나 치료 목적으로 검증된 것은 아니라고 명확히 밝혔다
Fig. 1: End-to-end pipeline and frozen evaluation benchmark. After segmentation QC, joint-label registration, registration/FOV QC, and same-structure R-local repair, the 383 corresponding meshes are divided into a development set used for model selection and a 76-case evaluation set withheld from every fit reported here (the graph architecture predates this partition; Section V). The orange block carries the conditional-mean equation; all comparators use the same split, panels, endpoint, and evaluation cases. CARE2026 and MM-WHS provide separate expert-surface validation without fitting or tuning.
Fig. 1: End-to-end pipeline and frozen evaluation benchmark. After segmentation QC, joint-label registration, registration/FOV QC, and same-structure R-local repair, the 383 corresponding meshes are divided into a development set used for model selection and a 76-case evaluation set withheld from every fit reported here (the graph architecture predates this partition; Section V). The orange block carries the conditional-mean equation; all comparators use the same split, panels, endpoint, and evaluation cases. CARE2026 and MM-WHS provide separate expert-surface validation without fitting or tuning.
TABLE I: Closest multistructure CT geometry resources. “Public package” means a direct public download of a reusable model or mesh cohort, assessed from the cited publications and their linked repositories at the time of writing; boundary tags or cut rings are not counted as separate anatomical surface bodies. None of the prior resources supports missing-label conditioning; this work conditions on any observed structure-block subset.
ResourceAnatomy outside the four chamber wallsPublic package
Hoogendoorn et al. [20]Detailed multiregion whole-heart atlas, including great vesselsAtlas and SSM reported; no reusable model package located
Rodero et al. [2]AO and PA walls; PV and caval openings tagged; LAA body omittedSSM parameters and 1 000 synthetic finite-element meshes
Strocchi et al. [3]AO and PA walls; LAA, PV, SVC and IVC at cut rings24 patient-specific finite-element meshes
This workAO, PA, LAA, PV, SVC and IVC as separate surface blocks in one shared topologyTemplate, per-case displacement fields and reconstruction code
Fig. 2: The eleven-structure whole-heart template mesh (11 571 vertices) in anterior, posterior, left, and right views, each structure a colour-coded surface in fixed vertex correspondence. This single template is warped into every patient’s space to build the SSM.
Fig. 2: The eleven-structure whole-heart template mesh (11 571 vertices) in anterior, posterior, left, and right views, each structure a colour-coded surface in fixed vertex correspondence. This single template is warped into every patient’s space to build the SSM.
TABLE II: The three cohorts and the exclusions that produce them. The released internal metadata supplies age, sex, manufacturer, scanner model, tube voltage, pathology class, and study type; “n/a” attributes are absent from the released metadata and are not estimated, and “none” means no fitting or tuning. The 76-case evaluation list is frozen (split file released with the artifacts). Both CARE2026 exclusions fall in subset B; MM-WHS has no exclusions.
Internal (TotalSegmentator)CARE2026MM-WHS CT
RoleSSM fit; internal eval.indep. validationindep. secondary
Source cases63160 (A/B/G)20
Segmentation QC−206 (vol., extent, centr.) → 425
Affine check−2 (non-orthon.)
Registration QC−26 (>25 mm) → 399
FOV truncation−16
Retained38358 (20/18/20)20
Split/tuning307 dev. (5 folds) / 76 withheld†nonenone
Labelssilver, all 11silver 11, expert 7expert 7
Scanner models11 distinctn/an/a
Median age64 yearsn/an/a
Study mix336 non-cardiac, 47 cardiacn/an/a
Contrast phasen/an/an/a
ECG gating, phasen/a, likely mixedn/an/a
Spacing analysed1 mm isotropic1 mm isotropicper case → 1 mm
†The graph architecture predates the partition and may have seen
cases now in the 76-case evaluation list.
Fig. 3: The completion operator. A selection matrix 𝐏o extracts whichever structure blocks or partial vertex sets are observed. The covariance is fitted once; inference extracts the corresponding covariance blocks and evaluates the conditional mean. Observed coordinates are copied to the output; missing coordinates are predicted. The figure writes the operator in covariance form for readability; the implementation evaluates the low-rank coefficient-space posterior of Eq. (1), whose ridge acts on the posterior precision rather than on 𝐂o​o.
Fig. 3: The completion operator. A selection matrix 𝐏o extracts whichever structure blocks or partial vertex sets are observed. The covariance is fitted once; inference extracts the corresponding covariance blocks and evaluates the conditional mean. Observed coordinates are copied to the output; missing coordinates are predicted. The figure writes the operator in covariance form for readability; the implementation evaluates the low-rank coefficient-space posterior of Eq. (1), whose ridge acts on the posterior precision rather than on 𝐂o​o.
TABLE III: Eleven cardiac structures and template vertex counts. MYO is the left-ventricular wall; the other ten are blood-pool or luminal surfaces. The model therefore carries no right-ventricular or atrial wall and does not represent total myocardial mass.
IDStructureAbbrev.Vertices
1Left ventricleLV1 502
2MyocardiumMYO2 502
3Right ventricleRV2 002
4Left atriumLA1 002
5Right atriumRA1 002
6AortaAO1 202
7Pulmonary arteryPA802
8Left atrial appendageLAA402
9Pulmonary veinsPV602
10Superior vena cavaSVC402
11Inferior vena cavaIVC151
Total11 571
Fig. 4: Completion across six observed-structure panels, one internal TotalSegmentator case (TS, s0365) beside one external CARE2026 case (A_Case1020), each the median case of its cohort by the displayed Cond-G estimator’s mean missing-vertex error over the six panels. Observed structures are grey; the conditional-Gaussian model completes the rest, coloured by per-vertex error (mm). The estimator is the selected configuration of Table V, fitted on the 307 development cases in the R-local domain.
Fig. 4: Completion across six observed-structure panels, one internal TotalSegmentator case (TS, s0365) beside one external CARE2026 case (A_Case1020), each the median case of its cohort by the displayed Cond-G estimator’s mean missing-vertex error over the six panels. Observed structures are grey; the conditional-Gaussian model completes the rest, coloured by per-vertex error (mm). The estimator is the selected configuration of Table V, fitted on the 307 development cases in the R-local domain.
TABLE IV: Per-structure warp-overlap Dice (SyN-warped patient segmentations vs template), CARE2026 (n=58): a registration-consistency measure, not accuracy against a manual reference. Structures marked † (LAA, PV, SVC, IVC) are silver-only. Chambers is the mean of the five per-chamber medians. p25 is the 25th percentile.
StructureMedianp25StructureMedianp25
LV0.9810.979AO0.8240.787
MYO0.9620.944PA0.9690.910
RV0.9740.965LAA†0.9410.938
LA0.9780.977PV†0.9180.891
RA0.9750.971SVC†0.7250.619
IVC†0.5280.461
Chambers: 0.974All 11: 0.889
TABLE V: Frozen 76-case internal evaluation completion MPVED (mm). Each k entry is the case mean over its fixed panels; the four-k column equally weights the four k values, with a case-bootstrap 95% CI. PPCA is the selected K=1,q=160 limit; local Cond-G reuses the exact global prediction by construction.
RepresentationMethodk=1k=3k=5k=9Four-k [95% CI]
RawMean shape11.21811.27811.18910.90911.149 [10.310, 12.040]
Nearest neighbour10.0969.0318.5088.6449.070 [8.621, 9.563]
Cond-G5.4793.9183.1933.0033.898 [3.629, 4.206]
β-VAE6.5965.2384.6934.7485.319 [4.991, 5.701]
PPCA (K=1)5.5243.9583.2483.0643.948 [3.683, 4.251]
Local Cond-G†5.4793.9183.1933.0033.898 [3.629, 4.206]
R-localMean shape11.15711.21511.13010.85811.090 [10.245, 11.990]
Nearest neighbour9.9428.9048.3528.5278.931 [8.503, 9.405]
Cond-G5.3043.7323.0072.8253.717 [3.480, 3.984]
β-VAE6.5225.1724.6214.6775.248 [4.935, 5.613]
PPCA (K=1)5.3483.7673.0612.8823.764 [3.530, 4.027]
Local Cond-G†5.3043.7323.0072.8253.717 [3.480, 3.984]
TABLE VI: Frozen paired family. Differences are challenger minus Cond-G, so positive values favour Cond-G. The R-local four-k primary is unadjusted; the next seven rows are the complete Holm family (m=7). The final row is an identity check rather than empirical evidence.
ContrastDifference (mm)95% CI (mm)Multiplicity result
R-local β-VAE minus Cond-G, four-k primary1.531[1.384, 1.711]Unadjusted p<0.001 (floor)
Raw β-VAE minus Cond-G, four-k1.421[1.264, 1.610]pHolm<0.002
R-local β-VAE minus Cond-G, k=11.218[1.023, 1.439]pHolm<0.002
R-local β-VAE minus Cond-G, k=31.440[1.298, 1.602]pHolm<0.002
R-local β-VAE minus Cond-G, k=51.613[1.466, 1.800]pHolm<0.002
R-local β-VAE minus Cond-G, k=91.852[1.689, 2.047]pHolm<0.002
R-local PPCA (K=1) minus Cond-G, four-k0.047[0.038, 0.056]pHolm<0.002
R-local local minus global Cond-G, four-k0.000[0.000, 0.000]Structural identity
TABLE VII: Additional non-linear baselines, R-local representation, MPVED (mm) with case-bootstrap 95% intervals. The internal column is the frozen 76-case split at four observed-structure counts; the MM-WHS column is exhaustive panels at k=1,3,5 over the seven expert-overlap structures, so only the ordering within a column is comparable. Both arms lie outside the prespecified Holm family and are descriptive; the latent-optimisation setting is not retuned externally. Cond-G, the β-VAE and the floors repeat Table V for reference; the same arms on CARE2026 expert surfaces, a different endpoint, are in Table VIII.
MethodTypeInternal 76MM-WHS (n=20)
Cond-GClosed-form linear3.717 [3.480, 3.984]4.823 [4.446, 5.226]
Latent optimisation‡Published method5.044 [4.746, 5.389]6.248 [5.755, 6.777]
β-VAEFeed-forward deep5.248 [4.935, 5.613]6.616 [6.109, 7.152]
SpiralNet++§Published operator6.710 [6.283, 7.218]8.916 [8.105, 9.741]
Nearest neighbourRetrieval floor8.931 [8.503, 9.405]10.907 [10.145, 11.684]
Mean shapePopulation floor11.090 [10.245, 11.990]13.174 [12.045, 14.419]
TABLE VIII: Additional non-linear baselines on CARE2026 expert surfaces, R-local, biventricular panel, mean ASSD / HD95 / Chamfer RMS (mm) over 58 cases. The endpoint is distance to expert manual segmentations and is not comparable to Table VII. AO and PA are suppressed for every method by the prespecified reference-closeness rule. Cond-G and the β-VAE repeat Table IX for reference.
MethodTypeLARA
Cond-GClosed-form linear3.275 / 7.846 / 4.1383.479 / 9.007 / 4.586
Latent optimisationPublished method3.902 / 8.941 / 4.8224.243 / 10.265 / 5.413
β-VAEFeed-forward deep3.934 / 8.990 / 4.8604.333 / 10.372 / 5.495
SpiralNet++Published operator4.299 / 9.701 / 5.2664.605 / 10.868 / 5.779
TABLE IX: R-local CARE biventricular native prediction-versus-expert surface metrics, pooled (n=58) and by acquisition subset (A/B/G contain 20/18/20 cases). The panel observes LV/MYO/RV; values are mean ASSD / HD95 / Chamfer RMS in mm. AO and PA are omitted after failing the reference-closeness rule.
PopulationMethodLARA
PooledCond-G3.275 / 7.846 / 4.1383.479 / 9.007 / 4.586
Pooledβ-VAE3.934 / 8.990 / 4.8604.333 / 10.372 / 5.495
ACond-G2.634 / 6.643 / 3.3533.035 / 8.085 / 3.989
Aβ-VAE3.103 / 7.553 / 3.9113.780 / 9.373 / 4.811
BCond-G4.010 / 9.642 / 5.0324.028 / 10.598 / 5.351
Bβ-VAE4.878 / 11.412 / 6.0344.706 / 11.572 / 6.011
GCond-G3.253 / 7.432 / 4.1193.427 / 8.496 / 4.494
Gβ-VAE3.915 / 8.247 / 4.7524.549 / 10.290 / 5.715
TABLE X: R-local CARE leave-one-expert-structure-out completion, pooled (n=58) and by acquisition subset (A/B/G contain 20/18/20 cases). Values are mean ASSD / HD95 / Chamfer RMS in mm. AO and PA are omitted because their references are too far.
Missing structurePopulationCond-Gβ-VAE
LVPooled2.362 / 6.407 / 3.1193.899 / 9.623 / 4.942
A1.818 / 4.744 / 2.3583.367 / 8.177 / 4.254
B3.697 / 10.361 / 4.9694.841 / 12.575 / 6.258
G1.704 / 4.510 / 2.2153.582 / 8.411 / 4.446
MYOPooled2.056 / 5.667 / 2.7292.943 / 7.583 / 3.769
A1.695 / 4.541 / 2.2122.526 / 6.257 / 3.171
B2.947 / 8.627 / 4.0363.847 / 10.332 / 5.046
G1.616 / 4.129 / 2.0682.547 / 6.434 / 3.219
RVPooled2.717 / 8.098 / 3.7703.377 / 8.555 / 4.335
A2.868 / 8.349 / 3.9693.016 / 8.073 / 3.994
B3.247 / 10.314 / 4.6194.148 / 10.495 / 5.305
G2.088 / 5.852 / 2.8073.045 / 7.290 / 3.801
LAPooled2.417 / 6.228 / 3.1923.393 / 7.837 / 4.207
A2.169 / 5.919 / 2.8712.577 / 6.451 / 3.266
B3.023 / 7.642 / 3.9194.445 / 10.407 / 5.495
G2.120 / 5.264 / 2.8603.262 / 6.909 / 3.989
RAPooled2.800 / 7.434 / 3.7633.878 / 9.530 / 4.988
A2.577 / 7.013 / 3.4583.474 / 8.729 / 4.448
B3.246 / 9.006 / 4.4414.372 / 10.762 / 5.607
G2.621 / 6.439 / 3.4563.836 / 9.224 / 4.972
TABLE XI: The evidence behind each structure; structures with identical status share a row. “Reference close enough” means the registered reference lies within the fixed distance limits of the expert surface, so a native comparison is meaningful; the internal benchmark scores only structures that are completion targets there. “Cond-G better” is the lower descriptive mean error against the β-VAE at every k. AO’s reference is too far because the expert annotations cover a shorter aortic extent than the whole-aorta reference, not because of misregistration (Section V-D).
StructuresInternal benchmarkCARE2026 expert checkMM-WHS expert check
LV, MYO, RVnever a completion targetreference close enough (leave-one-out panel)expert label, but observed as input
LA, RAnever a completion targetreference close enough (both panels)reference close enough; Cond-G better
AOCond-G better at every kreference too farreference too far
PACond-G better at every kreference too farreference close enough under R-local; Cond-G better
LAA, PV, SVC, IVCCond-G better at every kno expert labelno expert label
TABLE XII: Morphometric retention from the biventricular panel: per-structure volume agreement between the completion and the case’s own registered mesh. MAE is mean absolute error; medAPE is median absolute percentage error; Δ is the paired difference in absolute error against the development-only regression on the three observed chamber volumes, with its 95% case-bootstrap interval (negative favours completion). Vessel and appendage rows are segment volumes under the fixed template crop.
Internal (n=76)CARE2026 (n=58)
StructureMAE (mL)medAPEΔ [95% CI]MAE (mL), medAPE
LA7.998.3%−6.84 [−9.43, −4.43]18.47, 18.5%
RA9.618.9%−7.29 [−10.36, −4.35]20.89, 25.7%
AO13.815.6%−7.99 [−12.75, −3.33]23.08, 13.1%
PA6.748.5%−4.51 [−6.35, −2.68]13.39, 17.5%
LAA0.9310.7%−0.49 [−0.71, −0.27]1.41, 14.8%
PV1.4113.1%−0.01 [−0.31, +0.28]1.55, 13.7%
SVC3.1812.3%−0.60 [−1.30, +0.21]5.52, 22.6%
TABLE XIII: Nominal 95% ellipsoid coverage by out-of-support stratum, R-local, 76 evaluation cases (64 typical, 12 atypical). Strata come from the frozen development 95th-percentile Mahalanobis threshold. Exploratory and post-hoc; brackets are descriptive case-bootstrap intervals.
Cond-Gβ-VAE
Observed paneltypicalatypicaltypicalatypical
LV1.0000.946 [0.868, 0.998]0.1280.105 [0.071, 0.138]
LV, MYO, RV1.0000.944 [0.878, 0.991]0.1290.140 [0.104, 0.174]
Five chambers1.0000.974 [0.953, 0.993]0.1410.139 [0.103, 0.175]
Nine structures1.0000.999 [0.996, 1.000]0.2050.194 [0.110, 0.281]

왜 중요한가

서로 다른 병원이나 연구에서 심장의 일부만 라벨링한 CT 데이터를 하나로 합쳐 대규모 분석을 하려는 연구자들에게, 복잡한 딥러닝 모델을 만들기 전에 훨씬 간단하고 계산이 빠른 통계적 방법이 오히려 더 정확할 수 있음을 보여준다. 또한 딥러닝 모델의 성능을 주장할 때 무엇과 비교했는지가 결과를 크게 좌우한다는 점을 실증적으로 지적한다.

이 논문의 용어

  • 통계적 형상 모델(SSM) · 여러 사람의 장기 모양을 같은 기준점(정점)으로 정렬해 평균과 변화 패턴을 통계적으로 표현한 3D 모델
  • 정점 대응(vertex correspondence) · 서로 다른 사람의 심장 표면을 이루는 점들을 같은 위치끼리 짝지어 비교 가능하게 만드는 것
  • 조건부 가우시안 추정기(Cond-G) · 일부 값을 알고 있을 때 정규분포 가정 하에 나머지 값을 수학 공식으로 바로 계산하는 예측 방법
  • 그래프 β-VAE · 3D 메쉬 데이터를 그래프 형태로 다루며 압축·복원을 학습하는 딥러닝 모델의 한 종류
  • MPVED(평균 정점당 오차) · 예측한 표면과 실제 표면의 각 점 사이 거리를 평균 낸 정확도 지표(mm 단위)

논문 원문 초록 (영문)

Public cardiac cohorts annotate different subsets of the heart, so shapes from separate sources cannot be pooled without shared correspondence. Among released cardiac shape resources, none we identified carries the atrial appendage, pulmonary veins, and caval stumps as separate blocks in one mesh. Completion benchmarks also compare deep models against a least-squares projection onto shape modes, not the conditional estimator the same fitted model implies. We release an eleven- structure cardiac computed-tomography (CT) statistical shape model, built from 383 automatically labelled cases in 11 571-vertex correspondence, and compare completion estimators under one frozen internal split and endpoint. On a 76-case internal list held out from fitting, a closed-form conditional-Gaussian estimator reconstructed the missing non-chamber structures at 3.717 mm mean per-vertex error, averaged equally over one, three, five, and nine observed structures. A five-refit mask-conditioned graph variational autoencoder reached 5.248 mm and nearest-neighbour retrieval 8.931 mm. The paired difference was 1.531 mm (95% confidence interval 1.384 to 1.711), and the ordering held in a raw-coordinate sensitivity arm. Expert manual labels exist for 58 external CT cases, but our registered reference is close enough to score only five structures. There the closed-form estimator again had lower average surface distance, 95th-percentile Hausdorff distance, and Chamfer error for both completed atria. On a second public benchmark of 20 cases the reference was close enough for three of four completed structures, and the same ordering held there. Four structures have no expert reference. The released model and its completion operator support cohort-unification research on aligned CT, not clinical use.

저자 · Matej Gazda, Jakub Gazda, Juraj Gazda, Peter Drotar

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Matej Gazda et al., arXiv:2608.19932, arxiv-nonexclusive