Connect with us

AI

Grok 4.5 Targets Lawyers. Canada Penalised One for Using Grok

Cursor and SpaceXAI’s Grok 4.5 pushes into legal work weeks after a Canadian tribunal logged a record cost order over Grok-fabricated case citations.

Published

on

Grok 4.5, the first AI model SpaceXAI and Cursor built together, launched on Wednesday with coding, finance and legal work as its target markets. Cursor priced the base model at $2 per million input tokens and $6 per million output tokens, undercutting both Anthropic’s Opus family and OpenAI’s flagship.

Earlier in June, the Law Society Tribunal in Ontario imposed $31,150 in costs on Shahryar Mazaheri, a suspended lawyer whose Grok-generated case citations were deemed fabricated in a real disciplinary proceeding. It is the largest cost order against a lawyer for AI-fabricated authorities in any Canadian court or tribunal to date.

What Grok 4.5 Is Built to Do

Cursor’s launch blog post published the technical details on Wednesday. The model is positioned for “difficult, long-running tasks that require creatively using tools to solve problems,” spanning software engineering, data science, finance and legal work. Unlike Cursor’s previous coding-specialist model Composer 2.5, Grok 4.5 was kept deliberately broader in scope during training. The launch is also the first deliverable since SpaceX closed its $60 billion acquisition of Cursor earlier this year.

Grok 4.5 is a mixture-of-experts architecture trained jointly by Cursor and SpaceXAI on “trillions of tokens of Cursor data,” the post said. That dataset captures user interactions with codebases and software tools, including how developer-agents interact with their environments. Training data also drew on “high-quality STEM tasks, research papers, and other knowledge work.”

Cursor added new safeguards reflecting the model’s cybersecurity capabilities, designed to “preserve legitimate security work, including finding and patching vulnerabilities, while restricting the workflows most likely to cause harm.” Cursor said the safeguards are also built to detect and block bad actors.

How Cursor and SpaceXAI Trained the Model

Cursor describes the training as a reinforcement-learning push inside “realistic environments spanning both software engineering and broader knowledge work.” The environments were engineered to be “difficult enough that even frontier models fail at them,” because, as the company put it, “existing tasks stop teaching them anything new” once a model improves past them. Many problems had to be invented from scratch rather than drawn from existing benchmarks.

Building those environments at scale required a distributed agent system. Engineers specified what a problem was and how to verify a solution; “large groups of agents construct, test, and refine each environment,” Cursor wrote. The approach replaces what “would have taken teams of hundreds of engineers months to build.” Each environment is then tested under automated supervision and reused across training runs.

Cursor used Composer 2.5 to accelerate progress on Grok 4.5, the post said. The two models are in different weight classes, with Composer 2.5 to remain offered and new Composer-line models to come.

Recursive model-to-model training, where one model builds the data used to train the next, is now part of the joint Cursor-SpaceXAI pipeline, the post said. Cursor positions Composer 2.5 as the accelerant behind Grok 4.5’s broader training scope, beyond software engineering into law and finance. The same distributed agent harness is being readied for additional Composer-class models going forward.

The Price Tag and Where It Runs

Grok 4.5 is sold by the token. The base model is priced at $2 per million input tokens and $6 per million output tokens, with a faster variant carrying $4 input and $18 output, per the launch post. Both numbers undercut Anthropic’s Opus 4.8, priced at $5 per million input tokens and $25 per million output tokens, according to Axios.

Model Input $/M tokens Output $/M tokens
Grok 4.5 base $2 $6
Grok 4.5 fast $4 $18
Claude Opus 4.8 $5 $25
GPT 5.6 Luna $1 $6

Rollout covers:

  • Cursor across desktop, web, iOS, CLI and SDK
  • Grok Build
  • The SpaceXAI console

Cursor’s individual and team plans include “significant usage” of Grok 4.5 inside their first-party model pools, with usage doubled for the first week. Grok 4.5 is not yet available in the EU, Axios reported. Musk, writing on X, placed the model in Opus territory: “It is an Opus-class model, but faster, more token-efficient and lower cost,” he posted.

Cursor’s launch post pointed to Grok 4.5’s gains on coding and agentic evaluations as the model’s differentiators against its more expensive rivals. The company also acknowledged that an earlier snapshot of the Cursor codebase had been “accidentally” included in Grok 4.5’s training data, biasing one internal benchmark, CursorBench, in the new model’s favour; that data has been removed for future models.

A Documented Failure in a Real Legal Proceeding

Grok 4.5’s pitch lands against a documented case in the very market it is targeting. Earlier in June, the Law Society Tribunal in Ontario imposed $31,150 in costs on Shahryar Mazaheri, a suspended lawyer whose Grok-generated case citations were deemed fabricated in a real disciplinary proceeding. It is the largest cost order against a lawyer for AI-fabricated authorities in any Canadian court or tribunal to date.

Mazaheri, representing himself on a motion to lift the interlocutory suspension of his licence, filed a factum, a supplementary factum and two affidavits “all produced with generative AI.” The tribunal found the material cited tribunal decisions that did not exist, referenced real cases that did not support his arguments, and misused procedural rules as substantive law. In a January report on the hearing, Canadian Lawyer Magazine described Mazaheri’s November 2025 admission that he relied on Grok specifically and “failed to verify the output with sufficient care.”

But unlike a human, a large language model does not appreciate nuance or exercise judgment or use a moral compass when writing. It just does calculations, and then puts words in the order dictated by statistical probabilities.

The tribunal found Mazaheri’s “irresponsible use of artificial intelligence” “an additional and significantly aggravating factor,” per the tribunal’s June ruling on fabricated citations. The previous Canadian record for AI-related cost awards was $17,550, set by Reddy v. Saroya, a 2026 Alberta Court of Appeal decision.

Access-to-justice group Courtready.ca counts 168 decisions across 51 Canadian courts or tribunals in which someone submitted fabricated case law. Eighty-one percent involved self-represented parties, with the remainder involving lawyers. Courtready co-founder Tom Macintosh Zheng, a Toronto-based lawyer, said the volume is accelerating, with seven decisions in 2024, 86 in 2025, and 39 in the first six months of 2026.

Why Legal and Finance Work Lead the Pitch

Cursor listed legal work alongside software engineering, data science and finance in the launch post, framing Grok 4.5 as ready for the multi-step, document-heavy tasks that define each. Bloomberg has reported the legal push is part of SpaceXAI’s broader effort to recruit Wall Street clients and bulk up revenue. Anthropic and OpenAI have spent the past year positioning their coding agents at financial firms and corporate legal teams, the markets SpaceXAI is now chasing.

The release is SpaceXAI’s first model since SpaceX went public earlier this summer, and it lands in a week when OpenAI said it would release GPT 5.6 widely after the Trump administration asked for a staggered rollout of new models. The administration also briefly imposed foreign-access restrictions on Anthropic’s most capable models; both firms have since received clearance to proceed.

Axios reported a less obvious wrinkle: Grok 4.5 was trained using the same compute capacity SpaceXAI is leasing to Anthropic and Google. As SpaceXAI’s own compute needs grow, the publication wrote, it may have to choose between reserving capacity for its own models and selling it to competitors, a constraint that ties into both Anthropic’s $900 billion private fundraising round and the compute it sells to other labs under SpaceX’s $29.4 billion Google cloud lease.

What Cursor Did Not Disclose

Cursor’s launch post did not publish the full benchmark numbers for Grok 4.5 against named competitive models. The release relied on Musk’s X post showing some wins and some losses against Opus 4.8 across agentic coding evaluations. Cursor explicitly noted that SWE-Bench Pro and Terminal-Bench figures shown in the chart came from “self-reported scores for third-party models.” The single internal benchmark where Grok 4.5 leads, CursorBench, was tilted in the new model’s favour by the “accidentally” included earlier snapshot of the Cursor codebase, the company disclosed.

The safeguards built around Grok 4.5 are described in general terms, not in specifics an enterprise compliance team could audit. Cursor’s stated goal is to “preserve legitimate security work, including finding and patching vulnerabilities, while restricting the workflows most likely to cause harm.” For the legal use case Grok 4.5 is now being sold into, no published tests show the model’s accuracy on filings, citations, or rules-based reasoning, the exact regime where Grok failed in Mazaheri’s motion materials.

Frequently Asked Questions

What was the lawyer’s role in the case and what did the tribunal find?

Shahryar Mazaheri represented himself on a motion to lift the interlocutory suspension of his licence; he had been suspended in November 2024 amid allegations of involvement in fraudulent mortgage transactions. The tribunal found that the Grok-generated material cited tribunal decisions that did not exist, referenced real cases that did not support his arguments, and misused procedural rules as substantive law.

What safeguards did Cursor add around Grok 4.5?

Cursor added safeguards to “detect and block bad actors” and to distinguish between legitimate security research and harmful workflows. The launch post did not publish technical details an enterprise compliance team could evaluate or audit on its own.

How does Grok 4.5 differ from Composer 2.5?

Composer 2.5 is a smaller coding-specialist model in Cursor’s pre-Grok 4.5 lineup. Grok 4.5 is in a different weight class and was trained on STEM, research and broader knowledge work in addition to coding. Cursor has said it will continue shipping Composer-class models alongside Grok 4.5.

Why is SpaceXAI’s legal and finance push happening now?

Bloomberg reported the legal push is part of SpaceXAI’s wider effort to recruit Wall Street clients and bulk up revenue. The launch is SpaceXAI’s first model since SpaceX went public earlier this summer, and it lands in a week when OpenAI is releasing GPT 5.6 widely.

How many fabricated-citation cases have been tracked in Canada?

Access-to-justice group Courtready.ca counts 168 decisions across 51 Canadian courts or tribunals in which someone submitted fabricated case law. Eighty-one percent involved self-represented parties, with the remainder involving lawyers.

Logan Pierce is a writer and web publisher with over seven years of experience covering consumer technology. He has published work on independent tech blogs and freelance bylines covering Android devices, privacy focused software, and budget gadgets. Logan founded Oton Technology to publish clear, no nonsense tech news and reviews based on real hands on testing. He has personally tested and reviewed dozens of mid range and budget Android phones, written extensively about app privacy, and built and managed multiple WordPress publications over the past decade. Logan holds a bachelor's degree in English and studied digital marketing at a certificate level.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending