Connect with us

AI

GIFT Turns Near-Miss CAD Programs Into Training Data

MIT’s GIFT method lifts image-to-CAD shape match 12% and cuts search compute 80% by training on a model’s own near-miss programs, not new human CAD files.

Published

on

An MIT-IBM method called GIFT converts 2D designs into 3D CAD programs with a 12% higher mean shape match, using 80% less search compute. The quieter result is how it gets there. The system keeps a vision model’s near-miss programs, scores them with a CAD kernel, and folds those errors back into training so it no longer waits for a fresh pile of human CAD files.

Lead author Giorgio Giannone, a DeCoDE Lab research affiliate at MIT and a principal research scientist on Red Hat’s AI Innovation Team, wants an engineer to point the method at a weak CAD model, set a compute budget, and let the loop run. The work was presented at the International Conference on Machine Learning after a March 28, 2026 arXiv posting, and MIT published the lab account on July 16, 2026.

CAD Files With History Are Still Scarce

The paper’s claim is blunt about the bottleneck. Image-to-CAD models do not mainly fail because the network is too small. They fail because few public sets pair a picture of a part with the editable program that built it. Meshes and B-Rep snapshots are easy to scrape. Construction history is not, and that history is what an engineer actually edits.

ABC holds about 1 million B-Reps with the build steps stripped out. DeepCAD, from 2021, offers about 178k sketch-and-extrude histories. Fusion 360 Gallery, also 2021, has about 8.6k human-authored sequences in the same narrow vocabulary. Fillets, chamfers, lofts, and Booleans that show up on real drawings are thin on the ground.

Anna Clare Doris and colleagues at MIT already tried to punch through that wall in 2025 with CAD-Coder, an open-source CAD-Coder vision-language model trained to emit CadQuery Python from an image. For that model they built GenCAD-Code: 163,671 image-code pairs, with 147,289 in train, 7,355 in test, and 9,027 in validation. GIFT does not replace that set. It treats it as a mine, then asks the already-tuned model to fail on it on purpose.

PUBLIC CAD SETS THE MODELS CAN ACTUALLY TRAIN ON

Dataset Year Size What you get
ABC 2019 1 million B-Reps, no build history
DeepCAD 2021 178k Sketch-and-extrude histories
Fusion 360 Gallery 2021 8.6k Human sketch-and-extrude sequences
GenCAD-Code 2025 163,671 Images paired with CadQuery scripts
Zero-to-CAD 2026 ~1 million Synthetic, readable CAD sequences

Companies still keep production Fusion and SolidWorks trees off the public internet. That is the drought GIFT is built around. Error traces from a model that already writes CAD code carry more signal than another round of waiting for donated files, because the mistakes cluster on the geometries the model cannot finish.

How GIFT Turns a Near-Miss Into a Lesson

GIFT stands for Geometric Inference Feedback Tuning. The bootstrapping image-to-CAD program synthesis paper starts from QwenVL-2.5-7B already tuned as CAD-Coder. For each training image it samples many candidate programs in parallel, executes them, and scores each solid against the ground-truth shape with intersection-over-union, or IoU, inside a CAD kernel.

Programs with IoU at or above 0.9 are treated as valid even when the code does not match the reference script. Programs between 0.5 and 0.9 are treated as structured failures. About 10% of samples fall below 0.5 and are thrown out as junk or non-executable. About 40% fall below 0.9, which is the band the authors actually want.

If we sample the model 10 times and it generates 10 correct answers to the same problem, then there is not much for it to learn. We care about the in-between cases, where the model might only solve the problem 50 percent of the time.

Giorgio Giannone, research affiliate, MIT DeCoDE Lab

Two mechanisms split that band. Soft-Rejection Sampling, called GIFT-REJECT in the MIT write-up, keeps diverse high-IoU programs as extra targets for the original image, so the model is not punished for a second valid way to build the same part. Failure-Driven Augmentation, GIFT-FAIL, takes a near-miss solid, renders it back into an image, and pairs that broken view with the original ground-truth code, so the next pass has to map a wrong-looking part onto the right program.

WHAT THE LOOP KEEPS AND WHAT IT DISCARDS

  • Kernel score: Each sampled program is executed and compared with the ground-truth solid by IoU.
  • GIFT-REJECT: Programs at IoU 0.9 or higher become extra code targets for the same image.
  • GIFT-FAIL: Programs between IoU 0.5 and 0.9 are rendered into new images and paired with the true code.
  • Dropped junk: Programs below IoU 0.5, including non-executable scripts, stay out of the training pool.

No person sits in that loop to patch syntax. The authors run the expensive search offline, then do ordinary supervised updates on the new pairs, which is how they say they capture test-time scaling without online reinforcement learning against a CPU CAD solver. Giannone’s line on the method is that they wanted data expansion informed by the model itself, not random crops and color jitter.

The Output Is Executable CadQuery Python

Consumer image-to-3D tools emit a mesh you can look at. GIFT emits a program an engineer can change. The vision-language model sees one image plus a short prompt, then writes CadQuery, a Python CAD library that builds solids on the Open CASCADE kernel and can export STEP, STL, and related CAD files.

That choice is the whole product. A tessellated mesh is painful to parameterize. A B-Rep without history is hard to edit. A CadQuery script is a short, named sequence of sketches, extrudes, cuts, and fillets, which is why the paper treats code as the latent program and the kernel as a deterministic decoder. Doris’s CAD-Coder pipeline already ran that path: image in, script out, kernel execute, IoU against the true solid.

CadQuery is built so you can write Python scripts that build parametric solids without a GUI, which is what you want if a model is going to emit the design. The script is the artifact that later hits a crash or durability study. If the code does not run, the 3D file never exists. Giannone notes that almost-correct CadQuery is easy for a standard vision-language model and fully executable, dimension-true code is the hard part.

What the 12% IoU Gain Measures

On the GenCAD-Code test distribution, the GIFT-trained model reaches a mean IoU of 0.782. The supervised CAD-Coder baseline they compare against sits at 0.698. That gap is the 12% relative lift in the abstract. The same GIFT run’s median IoU is 0.948, which means a fat tail of failures is still dragging the mean around. Shape match is better. It is not solved.

The 80% compute claim is a different measurement. After the extra data are baked into the weights, GIFT matches the peak pass@k of the supervised model while using 80% less inference-time search, so the user spends fewer samples at test time for the same ceiling. Single-shot accuracy and cheaper search are the two numbers MIT leaned on. They are not the same plot.

MEAN SHAPE MATCH ON SINGLE-VIEW IMAGE-TO-CAD

System Setting Mean IoU
CAD-Coder with GIFT Fine-tuned, single image 0.782
Supervised CAD-Coder Fine-tuned, single image 0.698
Gemini 3 Pro Zero-shot, high resolution 0.540
Gemini 3 Pro Zero-shot, low resolution 0.200

A Gemini 3 Pro zero-shot run in the paper’s appendix reaches 0.540 mean IoU at high resolution and 0.200 at low resolution, still well short of the specialist 7B model after GIFT. Other published image-to-CAD systems that add point clouds, depth, or extra views can post higher IoU on their own boards. GIFT’s bet is that a single image plus a kernel in the training loop is enough to stay in that race without a custom 3D decoder.

IoU here is not a beauty score. The kernel executes the predicted program and the ground-truth program, then measures how much of the two solids overlap. If that overlap is wrong, every virtual crash built on top of the file is wrong. That is why the authors started with geometry and left performance and manufacturability for later.

Red Hat and IBM Staffed the Loop

The author list is not a CAD vendor list. Giannone and Kai Xu are on Red Hat’s AI Innovation Team. Co-senior author Akash Srivastava is director of Core AI at IBM and a principal investigator at the MIT-IBM Computing Research Lab, which funded the work in part. Co-senior author Faez Ahmed leads MIT’s DeCoDE Lab in mechanical engineering and is also a lab PI. Amin Heyrani Nobari, an MIT postdoc, completes the DeCoDE side with Doris.

On July 6, 2026, Srivastava described GIFT as image-to-CAD program synthesis that uses a CAD kernel as the verifier to bootstrap training data, matching heavy test-time scaling with 80% less compute. That is the industrial reading: sell a loop that turns budgeted inference into better weights, sitting above whatever CAD seat the customer already pays for.

Ahmed’s case for the method is that most physical products start as CAD, that current image-to-CAD models still emit simple shapes, and that a model which learns from its own errors does not have to wait for more human-made data.

What excites me about this work is that it gives many image-to-CAD-code models a way to improve themselves, learning from their own errors rather than waiting for more human-made data.

Faez Ahmed, associate professor of mechanical engineering, MIT DeCoDE Lab

Giannone’s deployment picture matches that. Point GIFT at a weak model, set the budget, and let the mistakes become the next training set. IBM and Red Hat are in the business of that kind of loop. Autodesk and Dassault are in the business of the files the loop is trying not to need.

A Million CAD Programs With No Real Files

In April 2026 Autodesk Research posted a parallel answer to the same drought. Zero-to-CAD puts a language model inside a CAD environment with tools and docs, then generates, executes, and repairs programs until they run. The group released about one million executable CAD sequences, plus a 100,000-model subset picked for geometric diversity, covering Booleans, fillets, chamfers, lofts, sweeps, and shells rather than sketch-and-extrude alone.

Their paper on synthesizing CAD programs without real files says the quiet part out loud: convert compute into data. Fine-tunes on that synthetic pile, they report, reconstruct editable programs from multi-view images without real construction-history supervision. It is the vendor-side twin of GIFT. One team mines near-misses on a 163,671-pair academic set. The other samples a million parts from scratch because professional histories stay locked in proprietary kernels.

Neither path waits for a carmaker to donate its native CAD tree. Both still have to prove the resulting programs survive a drawing check, a tolerance stack, and a toolpath. GIFT’s own next sentence in the MIT account is that the authors want the same loop to teach performance and manufacturability, and to run on larger models and messier CAD tasks than GenCAD-Code.

The Factory Floor Still Waits

GIFT is a geometry tutor. Giannone said they started there because, in engineering problems, if the 3D shape is wrong, nothing else is correct. The paper does not claim the programs are easy to machine, cheap to tool, or valid against a drawing standard. Future work, as stated, is to expand the method so models write CAD that improves performance and manufacturability, then to try bigger backbones.

The mean IoU of 0.782, with a median of 0.948, is the honest snapshot. Many parts land close. The failures that pull the mean down are still in the set. A 7B specialist with a kernel in the loop beats a frontier model asked to improvise CadQuery from one image, and it does that without extra human labels. It does not walk a bracket from a sketch to a fixture.

For rapid prototyping, that is still a shorter path from a 2D view to a file you can drop into a solver. The second-order change is the training diet. The scarce object in this field was always someone else’s CAD history. GIFT, and Autodesk’s million synthetic programs, are two ways of refusing to wait for it. The first drawing that has to be made, inspected, and paid for is still ahead of both.

Frequently Asked Questions

When Was the GIFT Paper Posted?

The paper, titled GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric Feedback, was submitted to arXiv on March 28, 2026 as 2603.27448, then presented at the International Conference on Machine Learning. MIT’s public write-up followed on July 16, 2026, with funding credited in part to the MIT-IBM Computing Research Lab.

What File Formats Can GIFT’s CadQuery Code Produce?

CadQuery is an Apache 2.0 Python library on the Open CASCADE kernel. Once a GIFT program executes, the usual exports are STEP and DXF for CAD interchange, plus STL, VRML, AMF, and 3MF for mesh and print workflows, which is the step that turns the model’s text into a file a solver or a printer can open.

How Did CAD-Coder Score Before GIFT?

CAD-Coder, posted May 20, 2025 as arXiv 2505.14646 and later run at IDETC 2025, reported a 100% valid syntax rate on its image-to-CadQuery test and higher 3D solid similarity than GPT-4.5 and Qwen2.5-VL-72B. GIFT starts from that CAD-Coder checkpoint rather than from a general vision model.

How Many Programs Does GIFT Sample While Mining Data?

The authors scale candidate sampling across five budgets, N in {8, 16, 32, 64, 128}, and invert the temperature: small budgets explore with wider temperature, large budgets sample colder for precision. Those extra samples are for building the augmented set, not a requirement at every later user query once the weights have absorbed the search.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending