diff --git a/manuscript/DRAFT_ML4H_v2_full.md b/manuscript/DRAFT_ML4H_v2_full.md index cce8149..47e95a6 100644 --- a/manuscript/DRAFT_ML4H_v2_full.md +++ b/manuscript/DRAFT_ML4H_v2_full.md @@ -18,15 +18,15 @@ That a tumour's molecular phenotype can be predicted from haematoxylin-and-eosin ## 1. Introduction -Predicting molecular status from H&E histology is by now a mature field. Microsatellite instability, gene mutations and expression subtypes have been predicted with deep learning [Coudray 2018; Kather 2019, 2020; Naik 2020], and pathology foundation models have pushed this further into molecular subtypes and drug sensitivity [Fernandez-Romero 2026; Dawood 2024]. That these targets *can* be predicted is well established. +Research using AI to analyse histopathological H&E images has been pursued across several organs as digital pathology has spread [CITE-I1]. With the wider use of CLAM-family weakly supervised multiple-instance learning [CITE-I2], work expanded in urological cancers [CITE-I3], breast cancer [CITE-I4], pancreatic cancer [CITE-I5], and other settings, and knowledge distillation and pathology foundation models have improved performance [CITE-I6]. Within this field, there has been persistent interest in predicting the molecular state of tissue from images. The reason lies in what is being replaced. IHC staining and tissue-destructive molecular tests, the usual methods for assessing molecular state, are generally costly and slow, whereas H&E staining is relatively inexpensive and is already acquired in routine care [CITE-I7]. Yet these molecular tests play important roles in early detection, prognostic prediction, and treatment direction across several cancer types [CITE-I8]. If inexpensive images can substitute for expensive tests, the potential gain is large. And the basic fact that molecular state can be learned and predicted from H&E has been shown repeatedly [CITE-I9]. -But being predictable does not mean it is acceptable to replace a molecular test clinically. Reporting predictive performance alone is silent about the clinical cost of substitution — the loss incurred when a wrong prediction assigns the wrong treatment. The same AUROC carries entirely different clinical consequences depending on which treatment decision the error lands in. This gap is where the present work sits. +But being predictable does not mean it is acceptable to replace a molecular test clinically. Reporting predictive performance alone is silent about the clinical cost of substitution — the loss incurred when a wrong prediction assigns the wrong treatment [CITE-I10]. The same AUROC carries entirely different clinical consequences depending on which treatment decision the error lands in [CITE-I11]. This gap is where the present work sits. We propose a cost-of-substitution frame. By converting prediction errors into the misassignment cost of treatment routing, we ask, for each molecular axis, whether H&E can substitute cheaply or whether molecular testing is required. The criterion is safety of substitution, not predictability. The frame does not predict drug response; it operationalises only the substitution cost from marker to treatment assignment, and it takes no drug structure as input. We test this with a pre-registered morphological-correlate law across five cancers, anchored on breast (plus lung, colorectal, gastric and head and neck). The law states that H&E can cheaply stand in for a test only when the molecular alteration has a morphological correlate recognisable at H&E resolution. The five cancers are a deliberate boundary for testing the law, not an open pan-cancer atlas expansion; and sealing predictions before results does not by itself confer confirmatory strength — it provides claim discipline that suppresses post-hoc selection. -This paper makes four contributions. First, the cost-of-substitution frame itself, together with the separation of confirmable axes from undecided ones obtained by applying one pre-registered protocol across five cancers. Second, an honest negative anchor: the breast HER2 axis shows no signal supporting H&E-based substitution, and this negative is robust to H&E stain normalisation. Third, claim discipline — explicit adjudication of insufficient power on the many mutation and amplification axes that our pre-registered split cannot decide, rather than reporting only the axes that happen to score high. Fourth, the framing of a different question — "when is substitution safe?" — rather than a contest over predictive accuracy. Unlike single-cohort breast prediction [Fernandez-Romero 2026] or drug-sensitivity prediction [Dawood 2024], this study contributes a methodological frame that applies one pre-registered evaluation protocol and a substitution-cost lens across a multi-cancer cohort. An external treatment-outcome check (Yale pCR) and a spatial-transcriptomics mechanistic look are reported only as provisional, Critic-pending exploratory analyses (§R6, §R7), not as contributions. +This paper makes four contributions. First, the cost-of-substitution frame itself, together with the separation of confirmable axes from undecided ones obtained by applying one pre-registered protocol across five cancers. Second, an honest negative anchor: the breast HER2 axis shows no signal supporting H&E-based substitution, and this negative is robust to H&E stain normalisation. Third, claim discipline — explicit adjudication of insufficient power on the many mutation and amplification axes that our pre-registered split cannot decide, rather than reporting only the axes that happen to score high. Fourth, the framing of a different question — "when is substitution safe?" — rather than a contest over predictive accuracy. Unlike single-cohort breast prediction [CITE-I12] or drug-sensitivity prediction [CITE-I13], this study contributes a methodological frame that applies one pre-registered evaluation protocol and a substitution-cost lens across a multi-cancer cohort. An external treatment-outcome check (Yale pCR) and a spatial-transcriptomics mechanistic look are reported only as provisional, Critic-pending exploratory analyses (§R6, §R7), not as contributions. @@ -238,3 +238,77 @@ To test whether the anchor results are an artefact of uncorrected H&E stain vari - **Citations** are provisional (brackets) until machine-verified by `agents/critic/scripts/verify_citations.py`. - **Venue** — npj Precision Oncology vs ML4H 2026: format/length constraints ``; compression likely needed for a workshop venue (Leader decision). - **Reporting-standard mappings** (TRIPOD+AI done; CLAIM/PROBAST/STROBE pending) and **Table 1 (cohort characteristics)** to be attached as Supplement. + +--- + +## References (working) + +Markers `[CITE-Ix]` in the text resolve here. This section grows section by section; re-verify with `verify_citations.py` before submission. + +### Introduction + +Markers `[CITE-I1]`–`[CITE-I13]`. Every entry below was checked against the source or publisher page; nothing is entered from memory. + +**[CITE-I1]** Spread of digital pathology and computer-aided pathology +- Nam, S., Chong, Y., Jung, C. K., Kwak, T. Y., Lee, J. Y., Park, J., ... & Go, H. (2020). Introduction to digital pathology and computer-aided pathology. *Journal of Pathology and Translational Medicine, 54*(2), 125–134. + +**[CITE-I2]** Uptake of weakly supervised WSI learning and CLAM-family MIL +- Lu, M. Y., Williamson, D. F. K., Chen, T. Y., Chen, R. J., Barbieri, M., & Mahmood, F. (2021). Data-efficient and weakly supervised computational pathology on whole-slide images. *Nature Biomedical Engineering, 5*(6), 555–570. https://doi.org/10.1038/s41551-020-00682-w +- Ilse, M., Tomczak, J., & Welling, M. (2018). Attention-based deep multiple instance learning. *Proceedings of the 35th International Conference on Machine Learning (PMLR), 80*, 2127–2136. + +**[CITE-I3]** H&E AI studies in urological (prostate, bladder) cancer +- Paik, I., Lee, G., Lee, J., Kwak, T. Y., & Ha, H. K. (2025). Artificial intelligence–driven digital pathology in urological cancers: Current trends and future directions. *Prostate International*. +- Cho, Y., Shin, D., Hong, S., Lee, J., Park, S., Lee, G., ... & Ha, H. K. (2026). Efficient AI-driven multi-section whole slide image analysis for biochemical recurrence prediction in prostate cancer. *arXiv preprint* arXiv:2603.20273. https://arxiv.org/abs/2603.20273 + +**[CITE-I4]** H&E WSI AI studies in breast cancer +- Lee, G., Lee, J., Kwak, T. Y., Kim, S. W., Kwon, Y., Kim, C., & Chang, H. (2025). Assessing the risk of recurrence in early-stage breast cancer through H&E stained whole slide images. *Scientific Reports, 15*(1), 35069. https://doi.org/10.1038/s41598-025-16679-x +- Lee, J., Lee, G., Kwak, T. Y., Kim, S. W., Jin, M. S., Kim, C., & Chang, H. (2024). MurSS: A multi-resolution selective segmentation model for breast cancer. *Bioengineering, 11*(5), 463. +- Lee, G., Kim, C., Kwak, T. Y., Kim, S. W., & Chang, H. (2023). Predicting protein receptor status from H&E-stained images in breast cancer. *Cancer Research, 83*(7_Supplement), 5404. + +**[CITE-I5]** Extension to pancreatic and other organs +- Lee, J., Lee, G., Kwak, T. Y., Kim, S. W., & Chang, H. (2022). A deep learning based pancreatic adenocarcinoma survival prediction model applicable to adenocarcinoma of other organs. *Cancer Research, 82*(12_Supplement), 5060. + +**[CITE-I6]** Knowledge distillation and pathology foundation models improving performance +- Cho, Y., Lee, S., Lee, G., Lee, M., Park, J., & Shin, D. (2026). G2L: From giga-scale to cancer-specific large-scale pathology foundation models via knowledge distillation. *AAAI 2026 Workshop (W3PHIAI)* [oral]. https://arxiv.org/abs/2510.11176 +- Kim, H., Kwak, T. Y., Chang, H., Kim, S. W., & Kim, I. (2023). RCKD: Response-based cross-task knowledge distillation for pathological image analysis. *Bioengineering, 10*(11), 1279. +- Chen, R. J., Ding, T., Lu, M. Y., Williamson, D. F. K., Jaume, G., Song, A. H., ... & Mahmood, F. (2024). Towards a general-purpose foundation model for computational pathology. *Nature Medicine, 30*(3), 850–862. https://doi.org/10.1038/s41591-024-02857-3 + +**[CITE-I7]** Cost and turnaround burden of IHC and tissue-destructive molecular tests relative to H&E +- Erfani, P., Gaga, E., Hakizimana, E., Kayitare, E., Mugunga, J. C., Shyirambere, C., Milner, D. A., Shulman, L. N., Ruhangaza, D., & Fadelu, T. (2023). Breast cancer molecular diagnostics in Rwanda: A cost-minimization study of immunohistochemistry versus a novel GeneXpert mRNA expression assay. *Bulletin of the World Health Organization, 101*(1), 10–19. https://doi.org/10.2471/BLT.22.288800 +- Sharma, A., Shah, P., Ranade, M., Pai, T., Sahay, A., Patil, A., Shet, T., Gupta, H., Chauhan, D., Somal, P., Sancheti, S., & Desai, S. (2025). Digital pathology enabling lean management of HER2/neu testing in breast cancer. *Journal of Pathology Informatics, 19*, 100515. https://doi.org/10.1016/j.jpi.2025.100515 + +**[CITE-I8]** Clinical role of molecular tests in early detection, prognosis and treatment direction +- Zhou, Y., Tao, L., Qiu, J., Xu, J., Yang, X., Zhang, Y., Tian, X., Guan, X., Cen, X., & Zhao, Y. (2024). Tumor biomarkers for diagnosis, prognosis and targeted therapy. *Signal Transduction and Targeted Therapy, 9*, 132. https://doi.org/10.1038/s41392-024-01823-2 + +**[CITE-I9]** Repeated demonstrations that molecular state can be predicted from H&E +- Coudray, N., Ocampo, P. S., Sakellaropoulos, T., Narula, N., Snuderl, M., Fenyö, D., Moreira, A. L., Razavian, N., & Tsirigos, A. (2018). Classification and mutation prediction from non–small cell lung cancer histopathology images using deep learning. *Nature Medicine, 24*(10), 1559–1567. https://doi.org/10.1038/s41591-018-0177-5 +- Kather, J. N., Pearson, A. T., Halama, N., Jäger, D., Krause, J., Loosen, S. H., ... & Luedde, T. (2019). Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer. *Nature Medicine, 25*(7), 1054–1056. https://doi.org/10.1038/s41591-019-0462-y +- Kather, J. N., Heij, L. R., Grabsch, H. I., Loeffler, C., Echle, A., Muti, H. S., ... & Luedde, T. (2020). Pan-cancer image-based detection of clinically actionable genetic alterations. *Nature Cancer, 1*(8), 789–799. https://doi.org/10.1038/s43018-020-0087-6 +- Naik, N., Madani, A., Esteva, A., Keskar, N. S., Press, M. F., Ruderman, D., ... & Socher, R. (2020). Deep learning-enabled breast cancer hormonal receptor status determination from base-level H&E stains. *Nature Communications, 11*(1), 5727. https://doi.org/10.1038/s41467-020-19334-3 +- Schmauch, B., Romagnoni, A., Pronier, E., Saillard, C., Maillé, P., Calderaro, J., ... & Wainrib, G. (2020). A deep learning model to predict RNA-Seq expression of tumours from whole slide images. *Nature Communications, 11*(1), 3877. https://doi.org/10.1038/s41467-020-17678-4 + + +**[CITE-I10]** Clinical decision loss of substituting a molecular test — performance alone does not establish clinical acceptability +- Vickers, A. J., Van Calster, B., & Steyerberg, E. W. (2016). Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests. *BMJ, 352*, i6. https://doi.org/10.1136/bmj.i6 +- Vickers, A. J., & Elkin, E. B. (2006). Decision curve analysis: A novel method for evaluating prediction models. *Medical Decision Making, 26*(6), 565–574. https://doi.org/10.1177/0272989X06295361 +- Van Calster, B., Collins, G. S., Vickers, A. J., Wynants, L., Kerr, K. F., Barreñada, L., ... & Steyerberg, E. W. (2025). Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: Overview and guidance. *The Lancet Digital Health, 7*(12), 100916. + +**[CITE-I11]** Biomarkers guide different diagnostic, prognostic and targeted-treatment decisions, so the consequence of an error depends on the downstream decision +- Zhou, Y., Tao, L., Qiu, J., Xu, J., Yang, X., Zhang, Y., Tian, X., Guan, X., Cen, X., & Zhao, Y. (2024). Tumor biomarkers for diagnosis, prognosis and targeted therapy. *Signal Transduction and Targeted Therapy, 9*, 132. (= [CITE-I8]) +- Chakravarty, D., Gao, J., Phillips, S., Kundra, R., Zhang, H., Wang, J., ... & Schultz, N. (2017). OncoKB: A precision oncology knowledge base. *JCO Precision Oncology, 1*, 1–16. +- Griffith, M., Spies, N. C., Krysiak, K., McMichael, J. F., Coffman, A. C., Danos, A. M., ... & Griffith, O. L. (2017). CIViC is a community knowledgebase for expert crowdsourcing the clinical interpretation of variants in cancer. *Nature Genetics, 49*(2), 170–174. + +**[CITE-I12]** Prior single-cohort or breast-focused H&E studies predicting receptor status, subtype or biomarkers +- Tafavvoghi, M., Sildnes, A., Rakaee, M., Shvetsov, N., Bongo, L. A., Busund, L. T. R., & Møllersen, K. (2025). Deep learning-based classification of breast cancer molecular subtypes from H&E whole-slide images. *Journal of Pathology Informatics, 16*, 100410. +- Farahmand, S., Fernandez, A. I., Ahmed, F. S., Rimm, D. L., Chuang, J. H., Reisenbichler, E., & Zarringhalam, K. (2022). Deep learning trained on hematoxylin and eosin tumor region of interest predicts HER2 status and trastuzumab treatment response in HER2+ breast cancer. *Modern Pathology, 35*(1), 44–51. +- Gamble, P., Jaroensri, R., Wang, H., Tan, F., Moran, M., Brown, T., ... & Chen, P. H. C. (2021). Determining breast cancer biomarker status and associated morphological features using deep learning. *Communications Medicine, 1*(1), 14. +- Couture, H. D., Williams, L. A., Geradts, J., Nyante, S. J., Butler, E. N., Marron, J. S., ... & Niethammer, M. (2018). Image analysis with deep learning to predict breast cancer grade, ER status, histologic subtype, and intrinsic subtype. *npj Breast Cancer, 4*(1), 30. +- Naik, N., Madani, A., Esteva, A., Keskar, N. S., Press, M. F., Ruderman, D., ... & Socher, R. (2020). Deep learning-enabled breast cancer hormonal receptor status determination from base-level H&E stains. *Nature Communications, 11*(1), 5727. (= [CITE-I9]) +- Fernandez-Romero, J., Ramos-Berciano, P., Perez-Perez, M., Benavides, D., Robles-Frias, A., Garcia-Gutierrez, J., & Macias-Garcia, L. (2026). Domain generalisation challenges in breast cancer molecular classification using foundation models: A cross-cohort exploratory study. *Medical & Biological Engineering & Computing, 64*, 2321–2331. https://doi.org/10.1007/s11517-026-03590-4 — 프로젝트가 지목한 **최근접 스쿱** + +**[CITE-I13]** Prior histology-based work framing the task as drug-sensitivity prediction +- Dawood, M., Vu, Q. D., Young, L. S., Branson, K., Jones, L., Rajpoot, N., & Minhas, F. U. A. A. (2024). Cancer drug sensitivity prediction from routine histology images. *npj Precision Oncology, 8*(1), 5. + +**카운슬 판정 기록 (codex 집필 → agy 적대검토 → codex 반박 1회 → Claude 정리).** 초안의 I10–I20 표식 11개 중 7개를 삭제했다. 사유는 전부 동일 — **우리 논문 자신의 주장·설계·결과·기여에 인용을 붙인 것**이다. 특히 (a) 논지 문장 "But being predictable does not mean..." 에 선행연구를 걸면 4문단 뒤 기여 주장("다른 질문의 정립")과 자기모순이 된다. (b) 염색정규화·conformal 문헌을 기여 목록에 붙인 것은 인용 채우기였다. (c) 사전등록 근거로 leakage·site-batch 문헌을 든 것은 논거가 다르다. +남은 자리가 4개뿐인 것은 Introduction ¶2–¶5 가 대부분 우리 프레임 설명이기 때문이다. **인용 밀도는 Methods(현재 0개)와 Results(현재 2개)에서 확보해야 한다.** +✅ **Bibliography settled (2026-08-31).** Authors and bibliographic details are confirmed for every Introduction reference. What remains depends on publication progress: Paik 2025 has no volume/issue/pages assigned yet (Prostate International, PII S2287888225000066), and Cho 2026 (prostate) is an arXiv preprint with no final venue. Re-check at proof stage. G2L 2026 was corrected to an AAAI 2026 **workshop** (W3PHIAI, oral), not the main conference. diff --git a/manuscript/DRAFT_ML4H_v2_full_ko.md b/manuscript/DRAFT_ML4H_v2_full_ko.md index d992a29..efae498 100644 --- a/manuscript/DRAFT_ML4H_v2_full_ko.md +++ b/manuscript/DRAFT_ML4H_v2_full_ko.md @@ -20,15 +20,15 @@ ## 1. Introduction -H&E 조직 이미지에서 종양의 분자 상태를 예측하는 연구는 이미 성숙한 분야다. 미세위성 불안정성·유전자 변이·발현 아형이 딥러닝으로 예측되어 왔고[Coudray 2018; Kather 2019, 2020; Naik 2020], 병리 파운데이션 모델이 그 성능을 끌어올리면서 분자 아형과 약물 감수성으로까지 확장되었다[Fernandez-Romero 2026; Dawood 2024]. 이 표적들이 *예측된다*는 명제 자체는 널리 입증되었다. +조직병리 H&E 이미지를 AI로 분석하려는 연구는 디지털 병리의 확산과 함께 여러 장기에서 이루어져 왔다[CITE-I1]. CLAM 계열의 weakly-supervised multiple-instance learning이 퍼지면서[CITE-I2], 비뇨기암[CITE-I3]·유방암[CITE-I4]·췌장암[CITE-I5] 등에서 연구가 활발히 이루어졌고, 지식 증류와 병리 파운데이션 모델이 그 성능을 끌어올렸다[CITE-I6]. 그중에서도 이미지에서 조직의 분자 상태를 예측하려는 요구는 계속되어 왔다. 그 이유는 대체 대상 쪽에 있다. 분자 상태를 확인하는 통상적 방법인 IHC 염색이나 조직파괴적 분자검사는 대체로 비싸고 오래 걸리는 반면, H&E 염색은 상대적으로 저렴하고 통상 진료에서 이미 촬영된다[CITE-I7]. 그런데 이 분자검사들은 여러 암종에서 조기 발견·예후 예측·치료 방향 결정에 중요한 역할을 한다[CITE-I8]. 값싼 영상이 비싼 검사를 대신할 수 있다면 얻는 것이 크다는 뜻이다. 그리고 H&E로부터 분자 상태를 학습·예측할 수 있다는 것 자체는 반복적으로 입증되어 왔다[CITE-I9]. -그러나 예측된다는 것이 곧 분자검사를 임상적으로 대체해도 된다는 것을 뜻하지는 않는다. 예측 성능만 보고하는 관행은 대체가 초래하는 임상적 비용, 즉 잘못된 예측이 잘못된 치료를 배정할 때 발생하는 손실을 말하지 않는다. 같은 AUROC라도 그 오차가 어떤 치료 결정에서 발생하느냐에 따라 임상적 대가는 전혀 다르다. 이 간극이 이 논문의 자리다. +그러나 예측된다는 것이 곧 분자검사를 임상적으로 대체해도 된다는 것을 뜻하지는 않는다. 예측 성능만 보고하는 관행은 대체가 초래하는 임상적 비용, 즉 잘못된 예측이 잘못된 치료를 배정할 때 발생하는 손실을 말하지 않는다[CITE-I10]. 같은 AUROC라도 그 오차가 어떤 치료 결정에서 발생하느냐에 따라 임상적 대가는 전혀 다르다[CITE-I11]. 이 간극이 이 논문의 자리다. 우리는 cost-of-substitution 프레임을 제안한다. 예측 오류를 치료 라우팅의 오분류 비용으로 환산해, 각 분자 축에서 H&E가 값싸게 대체될 수 있는지 아니면 분자검사가 필수인지를 묻는다. 기준은 예측 가능성이 아니라 대체 안전성이다. 이 프레임은 약물 반응을 예측하지 않으며, 마커에서 치료 배정으로 가는 치환비용만 조작화하고, 약물 구조를 입력으로 받지 않는다. 이를 유방 앵커에 폐·대장·위·두경부를 더한 다섯 암종의 사전등록된 형태학적 상관물 법칙으로 검정한다. 법칙의 요지는, 어떤 분자 변이가 H&E 해상도에서 알아볼 수 있는 형태학적 상관물을 가질 때에만 H&E가 그 검사를 값싸게 대신할 수 있다는 것이다. 다섯 암종은 법칙을 검정하기 위한 의도된 경계이지 열린 pan-cancer 아틀라스 확장이 아니며, 예측을 결과 이전에 봉인하는 사전등록은 확증 강도 자체를 주는 것이 아니라 사후 선택을 억제하는 claim 규율을 제공한다. -이 논문의 기여는 넷이다. 첫째, 치환비용 프레임 그 자체와, 동일한 사전등록 규약 하나를 다섯 암종에 적용해 확증 가능한 축과 미결 축을 구분한 것이다. 둘째, 정직한 음성 앵커다 — 유방 HER2 축은 H&E 기반 대체를 지지하는 신호를 보이지 않으며, 이 음성은 H&E 염색 정규화에 견고하다. 셋째, claim 규율이다 — 우리의 사전등록 분할이 판정할 수 없는 다수의 변이·증폭 축에서 점수가 높게 나온 축만 보고하는 대신 검정력 부족을 명시적으로 판정한 것이다. 넷째, 예측 정확도 경쟁이 아니라 "언제 대체가 안전한가"라는 다른 질문의 정립이다. 유방 단일 코호트 예측[Fernandez-Romero 2026]이나 약물 감수성 예측[Dawood 2024]과 달리, 본 연구는 동일한 사전등록 평가 규약 하나와 치환비용 렌즈를 다암종 코호트에 적용하는 방법론적 틀을 기여한다. 외부 치료결과 점검(Yale pCR)과 공간전사체 기전 관찰은 기여가 아니라 잠정적·Critic 대기의 탐색적 분석(§R6, §R7)으로만 보고한다. +이 논문의 기여는 넷이다. 첫째, 치환비용 프레임 그 자체와, 동일한 사전등록 규약 하나를 다섯 암종에 적용해 확증 가능한 축과 미결 축을 구분한 것이다. 둘째, 정직한 음성 앵커다 — 유방 HER2 축은 H&E 기반 대체를 지지하는 신호를 보이지 않으며, 이 음성은 H&E 염색 정규화에 견고하다. 셋째, claim 규율이다 — 우리의 사전등록 분할이 판정할 수 없는 다수의 변이·증폭 축에서 점수가 높게 나온 축만 보고하는 대신 검정력 부족을 명시적으로 판정한 것이다. 넷째, 예측 정확도 경쟁이 아니라 "언제 대체가 안전한가"라는 다른 질문의 정립이다. 유방 단일 코호트 예측[CITE-I12]이나 약물 감수성 예측[CITE-I13]과 달리, 본 연구는 동일한 사전등록 평가 규약 하나와 치환비용 렌즈를 다암종 코호트에 적용하는 방법론적 틀을 기여한다. 외부 치료결과 점검(Yale pCR)과 공간전사체 기전 관찰은 기여가 아니라 잠정적·Critic 대기의 탐색적 분석(§R6, §R7)으로만 보고한다. @@ -240,3 +240,80 @@ CLAM-SB attention MIL을 사용하였다(hidden 512·attention 256, 40–50 epoc - **인용**은 `agents/critic/scripts/verify_citations.py`로 기계 검증하기 전까지 잠정(대괄호)이다. - **Venue** — npj Precision Oncology vs ML4H 2026: 형식/분량 제약 ``; 워크숍 venue에는 압축 필요 가능(Leader 결정). - **보고 표준 매핑**(TRIPOD+AI 완료; CLAIM/PROBAST/STROBE 대기) 및 **Table 1(코호트 특성)**을 Supplement로 첨부. + +--- + +## 참고문헌 (작업본) + +본문 표식 `[CITE-Ix]` 에 대응한다. 섹션별로 늘려 나가며, 최종 제출 시 `verify_citations.py` 로 전수 재검증한다. + +### Introduction + +본문 표식 `[CITE-I1]`–`[CITE-I13]` 에 대응한다. 아래 서지는 원문 또는 출판사 페이지에서 대조했으며 추정 기입은 없다. + +**[CITE-I1]** 디지털 병리·computer-aided pathology 의 확산 +- Nam, S., Chong, Y., Jung, C. K., Kwak, T. Y., Lee, J. Y., Park, J., ... & Go, H. (2020). Introduction to digital pathology and computer-aided pathology. *Journal of Pathology and Translational Medicine, 54*(2), 125–134. + +**[CITE-I2]** weakly-supervised WSI 학습과 CLAM 계열 MIL 의 확산 +- Lu, M. Y., Williamson, D. F. K., Chen, T. Y., Chen, R. J., Barbieri, M., & Mahmood, F. (2021). Data-efficient and weakly supervised computational pathology on whole-slide images. *Nature Biomedical Engineering, 5*(6), 555–570. https://doi.org/10.1038/s41551-020-00682-w +- Ilse, M., Tomczak, J., & Welling, M. (2018). Attention-based deep multiple instance learning. *Proceedings of the 35th International Conference on Machine Learning (PMLR), 80*, 2127–2136. + +**[CITE-I3]** 비뇨기암(전립선·방광) H&E AI 연구 +- Paik, I., Lee, G., Lee, J., Kwak, T. Y., & Ha, H. K. (2025). Artificial intelligence–driven digital pathology in urological cancers: Current trends and future directions. *Prostate International*. +- Cho, Y., Shin, D., Hong, S., Lee, J., Park, S., Lee, G., ... & Ha, H. K. (2026). Efficient AI-driven multi-section whole slide image analysis for biochemical recurrence prediction in prostate cancer. *arXiv preprint* arXiv:2603.20273. https://arxiv.org/abs/2603.20273 + +**[CITE-I4]** 유방암 H&E WSI AI 연구 +- Lee, G., Lee, J., Kwak, T. Y., Kim, S. W., Kwon, Y., Kim, C., & Chang, H. (2025). Assessing the risk of recurrence in early-stage breast cancer through H&E stained whole slide images. *Scientific Reports, 15*(1), 35069. https://doi.org/10.1038/s41598-025-16679-x +- Lee, J., Lee, G., Kwak, T. Y., Kim, S. W., Jin, M. S., Kim, C., & Chang, H. (2024). MurSS: A multi-resolution selective segmentation model for breast cancer. *Bioengineering, 11*(5), 463. +- Lee, G., Kim, C., Kwak, T. Y., Kim, S. W., & Chang, H. (2023). Predicting protein receptor status from H&E-stained images in breast cancer. *Cancer Research, 83*(7_Supplement), 5404. + +**[CITE-I5]** 췌장 등 타 장기로의 확장 +- Lee, J., Lee, G., Kwak, T. Y., Kim, S. W., & Chang, H. (2022). A deep learning based pancreatic adenocarcinoma survival prediction model applicable to adenocarcinoma of other organs. *Cancer Research, 82*(12_Supplement), 5060. + +**[CITE-I6]** 지식 증류·병리 파운데이션 모델이 성능을 끌어올림 +- Cho, Y., Lee, S., Lee, G., Lee, M., Park, J., & Shin, D. (2026). G2L: From giga-scale to cancer-specific large-scale pathology foundation models via knowledge distillation. *AAAI 2026 Workshop (W3PHIAI)* [oral]. https://arxiv.org/abs/2510.11176 +- Kim, H., Kwak, T. Y., Chang, H., Kim, S. W., & Kim, I. (2023). RCKD: Response-based cross-task knowledge distillation for pathological image analysis. *Bioengineering, 10*(11), 1279. +- Chen, R. J., Ding, T., Lu, M. Y., Williamson, D. F. K., Jaume, G., Song, A. H., ... & Mahmood, F. (2024). Towards a general-purpose foundation model for computational pathology. *Nature Medicine, 30*(3), 850–862. https://doi.org/10.1038/s41591-024-02857-3 + +**[CITE-I7]** IHC·조직파괴 분자검사의 비용·소요시간 부담 (H&E 대비) +- Erfani, P., Gaga, E., Hakizimana, E., Kayitare, E., Mugunga, J. C., Shyirambere, C., Milner, D. A., Shulman, L. N., Ruhangaza, D., & Fadelu, T. (2023). Breast cancer molecular diagnostics in Rwanda: A cost-minimization study of immunohistochemistry versus a novel GeneXpert mRNA expression assay. *Bulletin of the World Health Organization, 101*(1), 10–19. https://doi.org/10.2471/BLT.22.288800 +- Sharma, A., Shah, P., Ranade, M., Pai, T., Sahay, A., Patil, A., Shet, T., Gupta, H., Chauhan, D., Somal, P., Sancheti, S., & Desai, S. (2025). Digital pathology enabling lean management of HER2/neu testing in breast cancer. *Journal of Pathology Informatics, 19*, 100515. https://doi.org/10.1016/j.jpi.2025.100515 + +**[CITE-I8]** 분자검사의 조기 발견·예후·치료 방향 결정 역할 +- Zhou, Y., Tao, L., Qiu, J., Xu, J., Yang, X., Zhang, Y., Tian, X., Guan, X., Cen, X., & Zhao, Y. (2024). Tumor biomarkers for diagnosis, prognosis and targeted therapy. *Signal Transduction and Targeted Therapy, 9*, 132. https://doi.org/10.1038/s41392-024-01823-2 + +**[CITE-I9]** H&E 로부터 분자 상태 예측이 반복 입증됨 +- Coudray, N., Ocampo, P. S., Sakellaropoulos, T., Narula, N., Snuderl, M., Fenyö, D., Moreira, A. L., Razavian, N., & Tsirigos, A. (2018). Classification and mutation prediction from non–small cell lung cancer histopathology images using deep learning. *Nature Medicine, 24*(10), 1559–1567. https://doi.org/10.1038/s41591-018-0177-5 +- Kather, J. N., Pearson, A. T., Halama, N., Jäger, D., Krause, J., Loosen, S. H., ... & Luedde, T. (2019). Deep learning can predict microsatellite instability directly from histology in gastrointestinal cancer. *Nature Medicine, 25*(7), 1054–1056. https://doi.org/10.1038/s41591-019-0462-y +- Kather, J. N., Heij, L. R., Grabsch, H. I., Loeffler, C., Echle, A., Muti, H. S., ... & Luedde, T. (2020). Pan-cancer image-based detection of clinically actionable genetic alterations. *Nature Cancer, 1*(8), 789–799. https://doi.org/10.1038/s43018-020-0087-6 +- Naik, N., Madani, A., Esteva, A., Keskar, N. S., Press, M. F., Ruderman, D., ... & Socher, R. (2020). Deep learning-enabled breast cancer hormonal receptor status determination from base-level H&E stains. *Nature Communications, 11*(1), 5727. https://doi.org/10.1038/s41467-020-19334-3 +- Schmauch, B., Romagnoni, A., Pronier, E., Saillard, C., Maillé, P., Calderaro, J., ... & Wainrib, G. (2020). A deep learning model to predict RNA-Seq expression of tumours from whole slide images. *Nature Communications, 11*(1), 3877. https://doi.org/10.1038/s41467-020-17678-4 + + +**[CITE-I10]** 분자검사 대체의 임상 의사결정 손실 — 예측 성능만으로는 임상 수용 가능성이 서지 않는다 +- Vickers, A. J., Van Calster, B., & Steyerberg, E. W. (2016). Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests. *BMJ, 352*, i6. https://doi.org/10.1136/bmj.i6 +- Vickers, A. J., & Elkin, E. B. (2006). Decision curve analysis: A novel method for evaluating prediction models. *Medical Decision Making, 26*(6), 565–574. https://doi.org/10.1177/0272989X06295361 +- Van Calster, B., Collins, G. S., Vickers, A. J., Wynants, L., Kerr, K. F., Barreñada, L., ... & Steyerberg, E. W. (2025). Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: Overview and guidance. *The Lancet Digital Health, 7*(12), 100916. + +**[CITE-I11]** 바이오마커가 진단·예후·표적치료 결정을 각각 다르게 가르므로, 오류의 결과는 하류 결정에 달려 있다 +- Zhou, Y., Tao, L., Qiu, J., Xu, J., Yang, X., Zhang, Y., Tian, X., Guan, X., Cen, X., & Zhao, Y. (2024). Tumor biomarkers for diagnosis, prognosis and targeted therapy. *Signal Transduction and Targeted Therapy, 9*, 132. (= [CITE-I8]) +- Chakravarty, D., Gao, J., Phillips, S., Kundra, R., Zhang, H., Wang, J., ... & Schultz, N. (2017). OncoKB: A precision oncology knowledge base. *JCO Precision Oncology, 1*, 1–16. +- Griffith, M., Spies, N. C., Krysiak, K., McMichael, J. F., Coffman, A. C., Danos, A. M., ... & Griffith, O. L. (2017). CIViC is a community knowledgebase for expert crowdsourcing the clinical interpretation of variants in cancer. *Nature Genetics, 49*(2), 170–174. + +**[CITE-I12]** 선행 단일코호트·유방 중심 H&E 연구 (수용체·아형·바이오마커 예측) +- Tafavvoghi, M., Sildnes, A., Rakaee, M., Shvetsov, N., Bongo, L. A., Busund, L. T. R., & Møllersen, K. (2025). Deep learning-based classification of breast cancer molecular subtypes from H&E whole-slide images. *Journal of Pathology Informatics, 16*, 100410. +- Farahmand, S., Fernandez, A. I., Ahmed, F. S., Rimm, D. L., Chuang, J. H., Reisenbichler, E., & Zarringhalam, K. (2022). Deep learning trained on hematoxylin and eosin tumor region of interest predicts HER2 status and trastuzumab treatment response in HER2+ breast cancer. *Modern Pathology, 35*(1), 44–51. +- Gamble, P., Jaroensri, R., Wang, H., Tan, F., Moran, M., Brown, T., ... & Chen, P. H. C. (2021). Determining breast cancer biomarker status and associated morphological features using deep learning. *Communications Medicine, 1*(1), 14. +- Couture, H. D., Williams, L. A., Geradts, J., Nyante, S. J., Butler, E. N., Marron, J. S., ... & Niethammer, M. (2018). Image analysis with deep learning to predict breast cancer grade, ER status, histologic subtype, and intrinsic subtype. *npj Breast Cancer, 4*(1), 30. +- Naik, N., Madani, A., Esteva, A., Keskar, N. S., Press, M. F., Ruderman, D., ... & Socher, R. (2020). Deep learning-enabled breast cancer hormonal receptor status determination from base-level H&E stains. *Nature Communications, 11*(1), 5727. (= [CITE-I9]) +- Fernandez-Romero, J., Ramos-Berciano, P., Perez-Perez, M., Benavides, D., Robles-Frias, A., Garcia-Gutierrez, J., & Macias-Garcia, L. (2026). Domain generalisation challenges in breast cancer molecular classification using foundation models: A cross-cohort exploratory study. *Medical & Biological Engineering & Computing, 64*, 2321–2331. https://doi.org/10.1007/s11517-026-03590-4 — 프로젝트가 지목한 **최근접 스쿱** + +**[CITE-I13]** 선행 조직영상 기반 약물감수성 예측 +- Dawood, M., Vu, Q. D., Young, L. S., Branson, K., Jones, L., Rajpoot, N., & Minhas, F. U. A. A. (2024). Cancer drug sensitivity prediction from routine histology images. *npj Precision Oncology, 8*(1), 5. + +**카운슬 판정 기록 (codex 집필 → agy 적대검토 → codex 반박 1회 → Claude 정리).** 초안이 ¶2–¶5 에 단 마커 11개 중 7개를 삭제했다. 사유는 전부 동일 — **우리 논문 자신의 주장·설계·결과·기여에 인용을 붙인 것**이다. (a) 논지 문장 "그러나 예측된다는 것이 곧 …" 에 선행연구를 걸면 4문단 뒤 기여 주장("다른 질문의 정립")과 자기모순이 된다. (b) 염색정규화·conformal 문헌을 기여 목록에 붙인 것은 인용 채우기다. (c) 사전등록 근거로 leakage·site-batch 문헌을 든 것은 논거가 다르다. +남은 자리가 4개뿐인 것은 Introduction ¶2–¶5 가 대부분 우리 프레임 설명이기 때문이다. **인용 밀도는 Methods(현재 0개)와 Results(현재 2개)에서 확보한다.** + +**추가 확보 — ① 임상 의사결정 손실은 해결(I10, Vickers 계열 3편 신규 등재).** 남은 2종은 이번 Introduction 에서 해당 마커를 삭제해 당장은 불필요하나, 사전등록 근거나 검정력·다중성 주장을 본문에 다시 세울 경우 ② 사전등록·registered report 방법론 ③ 통계적 검정력·다중성 통제 문헌이 필요하다. +✅ **서지 확정(2026-08-31).** Introduction 인용 문헌의 저자·서지 확인 완료. 남은 것은 출판 진행에 따라 바뀌는 항목뿐이다 — `paik-2025` 는 권·호·페이지가 아직 부여되지 않았고(Prostate International, PII S2287888225000066), `cho-2026-prostate-br` 은 arXiv 프리프린트로 최종 게재처 미정이다. 교정 단계에서 다시 확인한다. `cho-2026-g2l` 은 AAAI **본회의가 아니라 2026 워크숍(W3PHIAI) 구두발표**로 정정했다. +`I7` 실측 근거: IHC 바이오마커 분석 **환자당 US\$67.33**(전체 진단비 \$138.29의 48.7%) · HER2 IHC 재검 평균 **TAT 15.65일**(관행 워크플로 기준). 본문에 수치를 넣을지는 주저자 판단. diff --git a/research/REFERENCE_LIST.md b/research/REFERENCE_LIST.md index 144226d..139c2b1 100644 --- a/research/REFERENCE_LIST.md +++ b/research/REFERENCE_LIST.md @@ -4,11 +4,11 @@ > 자동생성(paper-info.yaml 기준) + 갭(인용됐으나 미분석)은 §마지막. 최종갱신 2026-07-17. -## §Intro/Related — H&E→분자 예측(선행·스쿱) (phenotype-prediction, 11편) +## §Intro/Related — H&E→분자 예측(선행·스쿱) (phenotype-prediction, 20편) | 상태 | 문헌 | 연도 | venue | 제목 | |---|---|---|---|---| -| **DEEP** | tafavvoghi-2024-jpi | 2024 | Journal of Pathology Informa | Deep learning-based classification of breast cancer | +| **DEEP** | tafavvoghi-2024-jpi | 2025 | J Pathol Inform 16:100410 | Deep learning-based classification of breast cancer molecular subtypes from H&E whole-slide images ⚠️slug는 2024이나 게재는 2025 | | brief | shamai-2024-commsmed | 2024 | Communications Medicine | Clinical utility of receptor status prediction and m | | brief | farahmand-2022-modpathol | 2022 | Modern Pathology | Deep learning trained on H&E tumor ROIs predicts HER | | brief | gamble-2021-commsmed | 2021 | Communications Medicine | Determining breast cancer biomarker status and assoc | @@ -19,6 +19,31 @@ | brief | kather-2019-msi | 2019 | Nature Medicine | Deep learning can predict microsatellite instability | | brief | couture-2018-npjbc | 2018 | npj Breast Cancer | Image analysis with deep learning to predict breast | | brief | coudray-2018-natmed | 2018 | Nature Medicine | Classification and mutation prediction from non-smal | +| brief | paik-2025-urologic-dp | 2025 | Prostate International | AI-driven digital pathology in urological cancers: c | +| brief | lee-2025-brca-recurrence | 2025 | Scientific Reports | Assessing the risk of recurrence in early-stage brea | +| brief | lee-2024-murss | 2024 | Bioengineering | MurSS: A multi-resolution selective segmentation mod | +| brief | cho-2026-g2l | 2026 | AAAI (accepted) | G2L: From Giga-Scale to Cancer-Specific Large-Scale | +| brief | cho-2026-prostate-br | 2026 | arXiv 2603.20273 | Efficient AI-Driven Multi-Section WSI Analysis for B | +| brief | lee-2023-receptor-status | 2023 | Cancer Res 83(7_Suppl) AACR | Predicting protein receptor status from H&E-stained | +| brief | lee-2022-pdac-survival | 2022 | Cancer Res 82(12_Suppl) AACR | A deep learning based pancreatic adenocarcinoma surv | +| brief | nam-2020-digitalpath-intro | 2020 | J Pathol Transl Med | Introduction to digital pathology and computer-aided | +| brief | kim-2023-rckd | 2023 | Bioengineering | RCKD: Response-based cross-task knowledge distillati | + + +## §Intro — 임상 맥락: 분자검사의 비용·소요시간·역할 + 임상 효용 평가 (clinical-context, 6편) + +> 치환비용 논지의 전제(대체 대상이 비싸고 느리며 임상적으로 중요하다)를 뒷받침. 전부 DOI·PMID 대조 완료. + +| 상태 | 문헌 | 연도 | venue | 제목 | 식별자 | +|---|---|---|---|---|---| +| brief | erfani-2023-rwanda-ihc-cost | 2023 | Bull World Health Organ 101(1):10-19 | Breast cancer molecular diagnostics in Rwanda: a cost-minimization study of immunohistochemistry versus a novel GeneXpert mRNA expression assay | doi:10.2471/BLT.22.288800 · PMID 36593782 | +| brief | sharma-2025-her2-tat | 2025 | J Pathol Inform 19:100515 | Digital pathology enabling lean management of HER2/neu testing in breast cancer | doi:10.1016/j.jpi.2025.100515 · PMID 41070375 | +| brief | zhou-2024-tumor-biomarkers | 2024 | Signal Transduct Target Ther 9:132 | Tumor biomarkers for diagnosis, prognosis and targeted therapy | doi:10.1038/s41392-024-01823-2 · PMID 38763973 | +| brief | vickers-2016-netbenefit | 2016 | BMJ 352:i6 | Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests | doi:10.1136/bmj.i6 · PMID 26810254 | +| brief | vickers-2006-dca | 2006 | Med Decis Making 26(6):565-574 | Decision curve analysis: a novel method for evaluating prediction models | doi:10.1177/0272989X06295361 | +| brief | vancalster-2025-perfmeasures | 2025 | Lancet Digit Health 7(12):100916 | Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance | arXiv:2412.10288 | + +**실측 수치(본문 인용 시).** erfani: IHC 바이오마커 분석 환자당 US$67.33 (전체 진단비 $138.29의 48.7%). sharma: HER2 IHC 재검 케이스 평균 TAT 15.65일(관행) → 8.775일(디지털). zhou: 조기선별·진단·예후·재발감시·표적치료를 포괄한 리뷰. ## §Related/Paper B — H&E→약물·cell-line (morphology-drug, 9편) @@ -146,7 +171,7 @@ | 상태 | 문헌(확정 서지) | slug | 우리 논문에서 | |---|---|---|---| -| DEEP | **Fernandez-Romero 2026** — Domain generalisation…FM (Med Biol Eng Comput 64) | fernandez-romero-2026-domaingen | 최근접 스쿱(유방 subtype, 외부열화) → 치환프레임 pivot | +| DEEP | **Fernandez-Romero 2026** — Domain generalisation challenges in breast cancer molecular classification using foundation models: a cross-cohort exploratory study (Med Biol Eng Comput 64:2321-2331, doi:10.1007/s11517-026-03590-4) | fernandez-romero-2026-domaingen | 최근접 스쿱(유방 subtype, 외부열화) → 치환프레임 pivot | | DEEP | **Kaczmarzyk 2026 (MAKO)** — ROR-P 재발위험 예측 (npj Digital Med 9:149) | kaczmarzyk-2026-mako | "예측 포화" 근거(⚠️ subtype 아니라 ROR-P) | | brief | **Shulman 2026 (Path2Space)** — AI 공간전사체 (Cell 189, 교신 Ruppin) | shulman-2026-path2space | 반대방향(복원 vs 치환 audit) ⚠️문서엔 "Kaminski" 오기 | | DEEP★ | **Farahmand 2022** (Mod Pathol 35:44) | farahmand-2022-modpathol | **Yale 앵커 head-to-head 바 = trastuzumab반응 CV AUC 0.80** (HER2 CV0.90/외부0.81) |