Connect with us

AI

Grok 4.5 Takes Cursor’s Coding Model Into Legal Work

Cursor and SpaceXAI’s Grok 4.5 lists legal work as a target after Grok already produced a record Ontario costs order for fake citations.

Published

on

Cursor and SpaceXAI shipped Grok 4.5 on July 8, 2026 with legal work on the job list. It is the coding startup’s first model built for more than software engineering, trained jointly as a mixture-of-experts system on trillions of tokens of Cursor session data.

The pitch is that hard training environments taught it to check its work. An earlier Grok already failed that test in an Ontario licensing case, and the lawyer who filed the output was ordered to pay $31,150 in costs.

Grok 4.5 Is Cursor’s First Model Past Software

Cursor’s own launch post does not call Grok 4.5 a legal specialist. It calls the system the first model built past software engineering, then lists the jobs it should handle on a computer: software, data science, finance, legal work, and the rest of knowledge work.

SpaceXAI is the name on the joint training run. xAI, the Grok lab Elon Musk founded, became part of SpaceX in February 2026, and the combined AI unit now ships under that label. Grok 4.5 is the first model Cursor and SpaceXAI trained together after SpaceX agreed to buy Cursor’s parent, Anysphere.

The company says the model can handle long jobs that need tools, and that it solves multistep tasks in under half the steps of comparable frontier models. New cybersecurity safeguards were added to match what the model can do. Cursor did not publish a detailed threat model with that line.

It runs in Cursor on desktop, web, iOS, the CLI, and the SDK. Individual and team plans put it in the first-party model pool. That pool now also includes Grok 4.6, which Cursor posted about on August 12, 2026, so Grok 4.5 is no longer the newest joint weight on the picker.

Trillions of Cursor Tokens, Then Legal Work

The training story is a coding flywheel with a wider mix poured on top. Cursor trained Composer 2.5, shipped May 18, 2026, as a coding specialist. For Grok 4.5 it kept the data mix broader on purpose.

HOW GROK 4.5 WAS TRAINED

  • Cursor sessions: Training included trillions of tokens of Cursor data that capture how people and agents touch codebases and software tools.
  • Wider mix: The run also drew on STEM tasks, research papers, and other knowledge work so the model would hold up outside a repo.
  • Hard environments: Reinforcement learning used difficult problems in realistic setups spanning software engineering and broader knowledge work.
  • Agent factory: Engineers specified a problem and how a solution is checked, then large groups of agents built, tested, and refined each environment, using Composer 2.5 to speed the next model.
  • Safeguards: Cursor says it added new protections that reflect the model’s cybersecurity abilities, without spelling out the controls in the launch note.

Those environments, Cursor wrote, teach the model to investigate problems, use tools, recover from mistakes, and verify results. Many of the tasks were built so even frontier models fail them. That is a software-style check: the test harness either passes or it does not.

These environments teach the model to investigate problems, use tools, recover from mistakes, and verify results.

Cursor, Introducing Grok 4.5, July 8, 2026

A passing unit test is not the same check as a real citation. Nothing in the launch materials describes a court-filings corpus, a citator, or a CanLII-style holdout set. Legal work sits on the job list because the model is being sold as a computer agent for knowledge work, not because Cursor published a legal benchmark.

A $31,150 Ontario Bill for Unchecked Grok Output

The Grok brand already has a Canadian tribunal record, and it is not Grok 4.5’s record. Shahryar Mazaheri, a lawyer whose licence had been suspended on an interim basis, used an earlier Grok to research and draft motion materials, then filed them without a proper check.

THE MAZAHERI FILING, 2024 TO 2026

  1. November 12, 2024: The Law Society Tribunal suspends Mazaheri’s licence on an interim basis.
  2. October 22, 2025: He moves to cancel or vary that suspension, then files a second motion to exclude Law Society evidence and to have the panel recuse itself.
  3. Late 2025: The panel finds the materials riddled with decisions that do not exist, real cases that do not stand for the points cited, and Tribunal rules treated as if they were substantive law.
  4. June 12, 2026: In Mazaheri v Law Society of Ontario, 2026 ONLSTH 112, the panel ordered costs of $31,150 to the Law Society of Ontario.

Paul Aterman, writing for the panel, called the use of artificial intelligence irresponsible and a significantly aggravating factor. Both motions failed. The panel said the rest of the profession should not pay for this kind of behaviour. Mazaheri told the tribunal the errors were entirely his responsibility and that they came from over-reliance on generative tools, in particular Grok, while he was filing on his own.

Tom Macintosh Zheng, a Toronto lawyer who co-founded the Courtready tracker of hallucinated-citation cases, has called that figure the largest costs order of its kind issued by a Canadian court or tribunal. The launch of a Grok model that lists legal work as a target does not move the duty. The filer still has to check the case.

SpaceX’s $60 Billion Cursor Purchase

Grok 4.5 is the first public fruit of a much larger paper trade. On June 16, 2026, days after SpaceX’s public listing, the rocket company said it would buy Anysphere, Cursor’s parent, for $60 billion in stock. SpaceX completed that purchase on August 14, 2026 and folded the startup into the SpaceXAI unit.

The companies had already been training together. SpaceX said that spring the joint model would land in Cursor and in Grok Build, xAI’s coding agent. Cursor’s July 8 post is that model arriving, with legal work named in the same breath as finance and data science.

The commercial logic is a loop, not a law library. Cursor sessions feed training. The trained model ships back into the editor. The editor collects more sessions. If that loop holds, SpaceXAI gets a live stream of how agents actually work, which is the scarce input for long-running tools. Law and finance are the enterprise doors that loop is being walked through.

Cursor still sells a cheaper coding specialist. Composer 2.5 remains on the picker as a different weight class, and Cursor says it will keep shipping models at that size. People who spend their day in the editor already treat that split as the product. On launch day the live argument was not about factums. It was whether Composer 2.5 would still carry multi-file refactors, and when a Composer 3 would show up.

What Grok 4.5 Costs Inside Cursor

On-demand use is billed per million tokens. The standard Grok 4.5 checkpoint is $2 per million input tokens and $6 per million output tokens. A Fast variant is $4 input and $18 output. Cursor doubled included usage for the first week after launch, through July 21, 2026.

GROK 4.5 TOKEN PRICES

Model Input, per million tokens Output, per million tokens
Grok 4.5 $2 $6
Grok 4.5 Fast $4 $18
Composer 2.5 $0.50 $2.50

Composer 2.5 is the low-cost coding line, not a Grok 4.5 discount. Docs list Grok 4.5’s provider as Cursor, with a 256k context window, agent thinking, and three effort levels (high is the default, with medium and low for cheaper, faster passes). On Cursor’s India-only Start plan, Grok 4.5 is fixed at medium effort unless the user is on Pro or higher.

Inside the editor the model can search files, hit the web, read and edit, run shell commands, drive a browser, and generate images. That tool belt is why Cursor can list legal work as a computer job. It is also why a bad citation can look finished. The agent can fetch, format, and hyperlink. It cannot accept professional liability.

OpenAI Sets a November 12 Cursor Shutoff

The purchase that produced Grok 4.5 also tripped a change-of-control clause. On August 28, 2026, OpenAI said it had told SpaceX it would wind down the OpenAI model contract that feeds Cursor, with a proposed shutoff date of November 12, 2026.

We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk’s companies violating contracts.

OpenAI, company statement, August 28, 2026

OpenAI said that after Musk acquired Twitter, now part of SpaceX, the company broke OpenAI’s contract terms, and that Musk admitted under oath earlier in 2026 that xAI, now also part of SpaceX, had violated those terms. The lab is holding the cutoff to the latest date the Cursor contract allows, while refusing future models, including its upcoming Astra system. It said it has worked with Cursor for nearly four years and that developers who rely on OpenAI inside the editor are the people who will feel the change.

So the first joint SpaceXAI-Cursor model is being asked to cover more than code, including legal work, while one of Cursor’s long-time model suppliers walks out. The verification lesson from Ontario has not been repealed. Grok 4.5 can list legal work as a task. The person who files the brief still owns the citation.

Frequently Asked Questions

Who Built Grok 4.5, Cursor or SpaceXAI?

Both. Cursor’s docs list the provider as Cursor and describe Grok 4.5 as a joint model with SpaceXAI, continued on trillions of tokens of Cursor data. SpaceXAI is the AI unit of SpaceX after xAI joined the rocket company in February 2026. The public launch lived in Cursor’s product surfaces first.

Does Grok 4.5 Replace Composer 2.5?

No. Cursor says they are different model weight classes, Composer 2.5 stays on sale, and new models of Composer’s size will keep coming. Cursor also pulled Grok 4.5 from its own CursorBench chart because an earlier snapshot of the Cursor codebase was accidentally included in training, a contamination it says it has removed for later models.

What Is Grok 4.5’s Context Window?

Cursor’s model docs set the context window at 256k tokens. Standard on-demand cache reads are billed at $0.50 per million tokens, and Fast cache reads at $1 per million. Those cache rates sit beside the $2/$6 and $4/$18 input-output list prices.

What Did the Ontario Tribunal Find About Grok?

The panel found Mazaheri’s AI-assisted motion materials incoherent, stuffed with fictitious authorities and with real cases used for the wrong propositions, then treated that failure as an aggravating factor on costs. He was allowed to refile after a case management conference. One cited style of error was a hyperlink for “Law Society of Ontario v. Mercer” that actually opened Law Society of Ontario v. Deokaran.

When Did SpaceX Close the Cursor Deal?

SpaceX and Cursor struck an April 2026 partnership that let SpaceX either buy Anysphere for $60 billion or pay $10 billion for the collaboration. SpaceX exercised the purchase option on June 16, 2026 in an all-stock deal and completed it on August 14, 2026. Grok 4.5, shipped July 8, 2026, was the first joint model released during that close.

Disclaimer: This article is news reporting and analysis of a product launch, a corporate purchase, and a published tribunal costs decision. It is informational only and is not legal advice, professional-conduct advice, or a recommendation to use any model in client work or court filings. Anyone considering AI tools for legal research or drafting should consult a licensed lawyer and the rules of their own law society or regulator before acting. Prices, model names, contract dates, and case statuses come from company posts, product docs, and tribunal reasons, and they can change as models and contracts move.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending