Connect with us

AI

Res-SwinFusion Posts 1.04 mm but Misses Many Landmarks

Res-SwinFusion records a 1.04 mm mean on ISBI skull X-rays, while 77.05% of harder-test landmarks fall inside 2 mm.

Published

on

Res-SwinFusion posted a 1.04 mm mean landmark error on the easier ISBI 2015 skull X-ray test. That mean falls under the 2 mm line many orthodontic papers treat as a clinical pass.

Hui Zhang and colleagues at West China Hospital, Sichuan University, put the hybrid model in the open BIO Integration study on August 17, 2026. On the harder 100-film test, only 77.05% of points landed inside 2 mm, so a clinic that trusts the mean still has a long tail to review.

A 2 mm Line Still Lets a Lot of Points Fail

Orthodontic software and research groups have used a 2 mm cutoff for years when they score automatic tracing. A point inside that radius is counted as a hit. A point inside 4 mm is often counted as merely acceptable. Those are different bars, and the tighter one is the one that shows up in success tables.

A mean under 2 mm can still hide a fat tail. Res-SwinFusion’s Test1 mean is 1.04 mm with a 0.90 mm standard deviation. Its 2 mm success rate on that same set is 88.56%, so about one point in nine still misses the line that clinics treat as correct.

Test2 is the stress set. The mean there is 1.37 mm with a 1.23 mm standard deviation, and the 2 mm hit rate drops to 77.05%. Just under one in four landmarks on that set would still need a hand correction if 2 mm is the rule the office uses.

even for well-trained orthodontists, it takes more than 0.5 h to manually locate all cephalometric landmarks.

Fulin Jiang and colleagues, Dentomaxillofacial Radiology

That time cost is why offices buy auto-trace tools. It is also why a 77.05% hit rate is not a small remainder. The missed points are the ones that eat the minutes the software was supposed to give back.

Res-SwinFusion Runs ResNet and Swin Branches Together

Zhang’s group, with Qinsheng Hu and Zekun Jiang as corresponding authors, built an encoder-decoder that keeps a convolutional branch and a Transformer branch in parallel. The CNN side is an ImageNet-pretrained ResNet, which holds local bone edges. The other side is a Swin Transformer in the base configuration, which is there for long-range layout across a full film.

Ze Liu’s group at Microsoft Research Asia introduced the shifted-window self-attention design so a vision Transformer could run on high-resolution images without paying a full global-attention bill. Res-SwinFusion uses that backbone because a downsampling CNN can drop the wide spatial cues that tell gonion from a similar curve elsewhere on the skull.

Two extra blocks sit between those branches and the heatmap decoder. The team released public training code and checkpoints with the paper, picking the run whose mean error sat closest to the average of ten repeats.

HOW THE HYBRID PUTS THE POINTS DOWN

  • Dual encoder: ResNet features and Swin-B features are extracted in parallel, then joined in a U-Net style decoder rather than a two-stage crop-and-refine pipeline.
  • Discrimination feature guidance: Pixel-level cues are passed across Swin blocks so similar curves, densities, and nearby tissue are less likely to swap identities.
  • Feature interactive aggregation: Global location features and local structure features are fused before heatmap decoding, replacing a plain skip connection.
  • Heatmap head: Landmarks are scored as Gaussian maps, then cleaned by dropping weak pixels, keeping the largest blob, and averaging the hottest remaining region.

Xinguo Wang of Winland Intelligent Imaging in Chengdu is on the author list, so the work is not only a university methods paper. Ahmad Alenezi of Kuwait University is listed as well. The orthopedic and rehab groups at West China Hospital, plus Ya’an People’s Hospital, supplied the clinical side.

What the ISBI, Hand, and Pelvis Tests Showed

The skull numbers come from the ISBI 2015 Grand Challenge: 400 cephalograms at 1935 by 2400 pixels with 0.1 mm spacing, 19 landmarks each, and the mean of two specialists as ground truth. The official split is 150 training films, 150 on Test1, and 100 on Test2. The team resized those films to 768 by 768 with padding to hold aspect ratio.

The hand set has 910 films and 37 landmarks on joints and fingertips, split 610 for training and 300 for testing, resized to 512 by 512. Physical scale assumes a 50 mm wrist width. The pelvic set is internal: 326 postoperative coronal films after total hip arthroplasty at West China Hospital, ethics file 2023-1975, ten landmarks marked by a radiologist and an orthopedic surgeon, 200 films for training and 126 for testing.

MEAN ERROR VERSUS THE 2 MM HIT RATE

Test set Test films / points Mean error Success rate
ISBI 2015 Test1 150 / 19 1.04 mm 88.56% within 2 mm
ISBI 2015 Test2 100 / 19 1.37 mm 77.05% within 2 mm
Public hand X-ray 300 / 37 0.63 mm 96.24% within 2 mm
Internal pelvic THA 126 / 10 1.44 mm 96.52% within 4 mm

On Test1 the 2.5 mm, 3 mm, and 4 mm hit rates were 93.16%, 96.04%, and 98.67%. On Test2 those figures were 84.68%, 89.79%, and 95.26%. Against a Thaler heatmap baseline the team retrained under the same protocol, 2 mm success rose 1.19 points on Test1 and 1.94 points on Test2.

The hand mean of 0.63 mm did not beat Thaler’s 0.61 mm on that set. The 2 mm and 4 mm hit rates of 96.24% and 99.71% were the figures Zhang’s group called stronger than prior methods. Pelvic success is reported at 4 mm, not 2 mm, and that is the set that looks like implant planning rather than an orthodontic tracing drill.

Manual Tracing Still Takes More Than Half an Hour

Jiang’s stomatology group at the same university trained a two-stage convolutional system on 9,870 cephalograms from 20 clinics, with more than 100 orthodontists refining the marks. That system’s mean error was 0.94 mm on its own data, a different collection from ISBI 2015, so the two means are not a head-to-head.

The workload those systems aim at is the click-through. Jiang put the full landmarking job at more than half an hour even for experienced orthodontists. Recent orthodontic reviews still put a complete tracing, landmarks plus measurements, at up to 15 minutes when the film is traced in ordinary practice software.

Those two clocks measure different jobs. Half an hour is locate-every-point. Fifteen minutes is a full analysis in a digital package. Both are long enough that an auto-trace that still needs a pass over one in four Test2 points will not empty the inbox, it will change which clicks remain.

Ablation on Test1 showed why the hybrid is not a cosmetic extra. Adding the Swin branch to a ResNet-101 U-Net baseline cut mean error by 0.72 mm and lifted 2 mm success to 72.90%. Swapping a plain skip for the fusion block brought mean error to 1.16 mm and added 13.61 points of 2 mm success. Each piece moved the needle. None of them erased the tail.

99 Landmarks and a Dentist in the Loop

Cephalometric auto-trace is already a shipped product category. CEPHX, cleared in a January 31, 2024 Food and Drug Administration letter as K231396, is Class II software under 21 CFR 892.2050. It places 99 cephalometric points on a 2D skull film for dentists, on patients 14 and older, with WebCeph (K220903) listed as the predicate.

FROM THE 2015 CHALLENGE TO A CLEARED TRACER

  1. 2015: The ISBI cephalometric grand challenge releases 400 labeled skull films that still anchor academic tests, including this one.
  2. 2021: Microsoft Research Asia publishes Swin Transformer, the shifted-window backbone later used as Res-SwinFusion’s global branch.
  3. January 31, 2024: FDA issues the CEPHX letter, with labeling that a licensed dentist still interprets the output.
  4. August 17, 2026: Res-SwinFusion goes online in BIO Integration, with code, ISBI numbers, and an internal hip series.

The CEPHX file is blunt about who is on the hook. It says results still need a licensed dentist for diagnosis, planning, and simulation. That is the product pattern this research would have to join if it ever left GitHub: a mean error in the paper, a human still named on the label.

Res-SwinFusion is not that product. It is a methods paper with a public repo. The 1.04 mm Test1 mean is in the same band as other recent academic tracers. The second-order fact is that crossing a 2 mm mean no longer proves a clinic can skip the review step, because FDA-cleared tools already assume that step, and Zhang’s own Test2 table leaves just under one in four points outside the tight bar.

The Pelvis Cases Are Still Short of the Clinical Target

The hip films are postoperative total hip arthroplasty views, average 4000 by 3200 pixels at 0.1 mm spacing, resized to 768 by 768. The recorded diagnoses include 163 osteoarthritis cases, 54 fractures, and 108 cases of femoral head necrosis. That is the anatomy the introduction said plain CNNs lose, overlapping metal, distorted contours, and structures that share a similar curve.

Mean error on that internal set is 1.44 mm. Success at 4 mm is 96.52%. Against a classic U-Net, the authors recorded success-rate gains of 18.84, 17.96, 15.53, and 12.66 points at the thresholds they tabulated. They still would not call it done.

the localization performance of our approach indicated a margin some distance away from the clinical target.

Hui Zhang and colleagues, BIO Integration

WHAT THE PAPER STILL DOES NOT COVER

  • Other modalities: The highlights say further validation is required before transfer to CT, MRI, or other modalities.
  • Hard anatomy: Discrimination guidance exists because similar radians, sizes, and densities still confuse a network, and the hip series is the test of that claim.

  • Training diet: Models ran 30 epochs on Tesla V100 GPUs with almost no augmentation beyond resize, which keeps the experiment clean and leaves domain shift for later.
  • Device status: Nothing in the paper is a cleared surgical planner. A dentist-in-the-loop label is already the rule for the closest commercial analog.

The 1.04 mm skull mean will be the number that travels. The 77.05% Test2 hit rate, and the authors’ line on the hip films, are the numbers that decide whether a planner can stop clicking. Code is up. The remaining misses are still a person with a mouse.

Frequently Asked Questions

What Datasets Did Res-SwinFusion Use?

The skull test is the ISBI 2015 Grand Challenge, 400 films with 19 landmarks and 0.1 mm spacing, using the mean of two specialists as ground truth. The hand set has 910 films and 37 landmarks. The pelvic set is 326 de-identified postoperative hip films from West China Hospital, not a public challenge dump.

What Is the Difference Between Mean Error and Success Rate?

Mean radial error averages every point’s distance to the human mark, so a few large misses can be diluted. Success detection rate counts the share of points inside a fixed radius, which is why 1.37 mm on Test2 still pairs with only 77.05% inside 2 mm. The paper also lists 2.5 mm, 3 mm, and 4 mm success figures for the skull tests.

Has Res-SwinFusion Been Cleared as a Medical Device?

No. It is a research model with a public GitHub repo. Separate commercial tracers such as CEPHX are Class II devices under 21 CFR 892.2050, limited in that clearance to patients 14 and older, and labeled so a licensed dentist interprets the output.

Did the Authors Test CT or MRI?

No. The published highlights state that further validation is required before transfer to CT, MRI, or other modalities. All reported numbers are on 2D X-ray, with Gaussian heatmap settings of sigma 12.5 and alpha 40 chosen after a convergence study.

Disclaimer: This article is news reporting and analysis of a published research paper and related device filings. It is informational only and is not medical advice, a diagnostic opinion, or a recommendation to use, buy, or rely on any landmarking software for patient care. Readers should consult a licensed orthodontist, radiologist, or orthopedic surgeon, and follow device labeling, before acting on landmark coordinates or treatment plans. Figures, clearance statuses, and intended-use statements reflect the papers and filings cited and may change with later studies or regulatory actions.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending