Published in Research

How accurate is ChatGPT in predicting FTMH surgery outcomes?

This is editorially independent content
6 min read

Findings from a recent study published in Retina tested whether ChatGPT-5 could predict long-term anatomical and functional outcomes after full-thickness macular hole (FTMH) surgery.

Those predictions were then measured against those made by retina specialists as well as the actual results. On paper, the model held its own. But the catch is what was driving the score.

Give me some background first.

Some context: While pars plana vitrectomy (PPV) closes most full-thickness macular holes, estimating how much vision a patient will recover—and whether a given hole will actually close—still leans heavily on specialist judgment.

The appeal of a large language model (LLM) here is obvious: feed it a clinical summary and a scan, get a prognosis to support preoperative counseling.

However, whether it can do that reliably is the open question investigators’ sought to answer.

Now, talk about the study.

This was a retrospective study run at a single medical center. ChatGPT-5 was pitted against two senior retina specialists and against actual 12-month outcomes for 50 eyes of 50 patients who underwent PPV for FTMH between 2021 and 2024.

The setup: De-identified clinical summaries were entered into the LLM with a standardized prompt.

For each eye, the model:

  • Predicted whether vision would improve, stay stable, or worsen by 12 months
  • Estimated the final best-corrected visual acuity (BCVA)
  • Predicted whether the hole would close

The retina specialists made parallel predictions on the same anonymized material.

Who was included in the study?

All patients were ages 15+ (mean age: 66.2 years). Each had complete preoperative data, optical coherence tomography (OCT) imaging, and a full 12 months of follow-up.

Baseline vision was poor, averaging 20/100, and the mean minimal hole diameter was 441µm.

What the model saw: age, sex, refractive status, lens status, ocular history, symptom duration, preoperative BCVA, hole diameter, surgical details, and a single foveal OCT B-scan.

And the findings?

At 12 months, closure occurred in 44 of 50 eyes (88%), and mean BCVA improved from 20/100 (0.7 ± 0.4 logMAR) to 20/63 (0.5 ± 0.5 logMAR), P = 0.03.

Functionally, 35 eyes (70%) improved by at least two Early Treatment Diabetic Retinopathy Study (ETDRS) lines, while 8 (16%) stayed stable and 7 (14%) worsened.

On anatomical prediction, ChatGPT-5 scored 90% accuracy against 72% to 86% for the specialists. On functional outcomes it scored 66% against 42% to 44%.

So how accurate were the LLM’s predictions?

ChatGPT-5 predicted closure in 49 of 50 eyes.

It correctly identified all 44 eyes that ultimately closed but correctly recognized only 1 of the 6 that remained open, illustrating an optimism bias toward predicting successful outcomes.

Tell me more.

The pattern was similar for visual outcomes. ChatGPT-5 performed best when vision improved, correctly predicting 60% of those cases, but its accuracy fell to just 13% when vision remained stable and 0% when vision worsened.

Its mean BCVA prediction error was 11.4 ± 10.8 letters, with a median error of 8 letters. Roughly 60% of estimates landed within two lines of the true outcome.

Also worth noting: The model picked up real radiologic features of FTMH from the scans, including intraretinal cystoid spaces, elevated hole edges, and choroidal hypertransmission.

  • However, its detection of vitreomacular adhesion was less consistent.

Limitations?

Considering this was a single-center, retrospective analysis of 50 eyes, the usual caveats around generalizability apply.

More to the point: The model's accuracy was inflated by a systematic optimism bias, which makes the raw numbers look better than the clinical reasoning behind them.

The authors flagged that larger prospective studies are needed before anything like this reaches practice.

Expert opinion?

No outside commentary was included in the study.

The authors' read: “At first glance, ChatGPT-5 seemed as effective as, or better than, retinal specialists in predicting postoperative FTMH surgery outcomes,” they wrote. “However, this apparent advantage was mainly due to a consistent overestimation of positive results, indicating a bias toward predicting FTMH closure and visual improvement.”

Anything else?

The authors did credit the model's diagnostic potential. Despite no domain-specific training on ophthalmic imaging, ChatGPT-5 interpreted OCT morphology from a single foveal B-scan plus a structured summary—which they described as “an emerging capacity for multimodal reasoning when textual and visual data are integrated.”

Their bottom line: The outputs need cautious interpretation to avoid misleading confidence.

And lastly: the take home.

ChatGPT-5 can generate quantitative acuity predictions and read OCT images, but its confidence runs ahead of its judgment, especially for the eyes that don't close or don't improve.

For now, AI-generated prognoses belong in the supportive-information column. Counseling patients on expected recovery after macular hole repair should still rest on evidence-based guidance and (human) retinal specialist input.