Connect with us

AI

Wispr Flow Puts $280 Million Into Its Own Speech Model

Wispr Flow raised $280 million at a $2 billion valuation to fund Canto, a speech model for noise and mixed languages, as Google gives dictation away.

Published

on

Wispr Flow raised $280 million at a $2 billion valuation on August 17, 2026, and said the first and largest part of that money is going into speech accuracy. Menlo Ventures led the Series B, taking total capital to $361 million, while the San Francisco company previewed Canto, its first in-house speech model.

The round looks like a bigger dictation-app check. The spend plan is a model company trying to outrun a free keyboard.

Most of the $280 Million Goes Into Accuracy

Tanay Kothari, Wispr’s chief executive and co-founder, wrote that a year ago most people still opened a voice chat by explaining why they did not trust it. He says those talks have stopped. Users now tell him how much time they save, and what they want Flow to do next.

That shift raises the cost of a wrong word. When voice is how someone writes a client note or a work doc, one misspelled name pulls them back to the keys, and the thought they were holding is gone. Kothari put the bar in one line.

The whole reason to talk instead of type is to stay inside your own train of thought.

Tanay Kothari, CEO, Wispr Flow

Accuracy, in that frame, is not a nice extra. It is the thing that decides whether talking stays faster than typing. Humans already speak four to five times faster than they type, Menlo partner Matt Kraning wrote the same day, and Flow’s whole pitch is to keep that gap without a cleanup pass.

FLOW IN ONE SNAPSHOT

  • Words written: People have written more than 60 billion words with Flow.
  • Work use: Flow is used by people at almost all of the Fortune 500 and over 10,000 enterprises.
  • Reach: Menlo says Flow is used in 162 countries and over 100 languages.
  • Growth: Kraning wrote that revenue has grown over 30x year over year.

Kothari said most of the new round goes into research that closes the remaining gap, and into putting that work everywhere people already talk. Canto is the first public piece of that spend.

Canto Drops Noisy Word Errors to 5-10%

Alongside the funding, Wispr posted a preview of its first proprietary speech model, Canto. Kothari described it on X as a 2B-parameter model trained for loud rooms, windy streets, a toddler in the background, and a second language mixed into the first. That 2B-parameter count is the model’s size. It is not the company’s $2 billion price.

Most speech systems, he wrote, are trained on clean recordings in a quiet room, with a good mic and an accent the model has heard before. Flow’s users are in cars, on sidewalks, and a foot from someone else’s call. In those hardest conditions, with background noise, wind, heavy accents, or music, Kothari said error rates fall from more than 30% of words to somewhere between 5% and 10%.

WHAT CANTO CHANGES IN TESTING

Setting Before Canto Canto preview
Hard rooms (noise, wind, music, heavy accents) More than 30% of words wrong 5% to 10% of words wrong
Everyday dictation Edits still common 30% to 35% fewer dictations need an edit
Score Wispr chases Close enough to send Zero edit rate, text left untouched

Zero edit rate is Wispr’s internal score: the share of speech that comes back right the first time and needs nothing. Not close enough. Not quick to fix. Untouched. Canto is still a preview, so those cuts are a claim about a model that has not fully shipped.

Even the better everyday number leaves work on the table. A 30% to 35% drop in the share of dictations that still need edits means a lot of lines still get a pass with the keyboard. On September 4, 2026, Wispr’s own product notes also described a recent drop in accuracy tied to infrastructure strain, mixed older models, and language routing bugs for UK English and Swiss German. The company said it has since put everyone on the same current model.

Google Already Put Dictation in Gboard

The feature layer around Flow is getting cheaper. On May 12, 2026, Google announced Rambler, a Gemini-powered dictation mode inside Gboard that strips filler, follows mid-sentence corrections, and handles code-switching, including English to Hindi in one thought. Google framed it as reinventing the keyboard, and said it would not store voice recordings. The first wave was aimed at Galaxy and Pixel phones. By early September it was a Pixel 11 headline feature.

Microsoft’s SwiftKey began testing its own AI voice mode in an Android beta around September 10, 2026, including an offline path Rambler does not offer. Other apps have been pitching free, unlimited dictation and on-device Mac tools that advertise lower release-to-text latency than Flow. Dictation, in other words, is showing up as a checkbox.

THE FREE AND NEARLY FREE RIVALS

  • Gboard Rambler: Gemini dictation in the default Android keyboard, with multilingual code-switching, first on Pixel 11 and Galaxy phones.
  • SwiftKey AI voice: A September 2026 beta that cleans rambling speech and can run offline.
  • On-device Mac apps: Small local tools built on Apple’s speech stack now sell speed and privacy against a cloud app.
  • Willow: A competing dictation product that has been marketing free, unlimited voice writing as the reason to stop paying for Flow.

Kraning wrote down the objection himself. Dictation looks like a feature, and free, mostly-right voice has been on phones for a decade with almost no one living in it. His answer is that “free and mostly right” loses to “zero edits, everywhere” for people who talk for a living, and that acting on what you meant takes context no OS toggle has. Menlo is not doubling down on a microphone button. It is doubling down on the claim that the text box is the next interface to disappear.

Google is also folding voice tools inside Gmail and Docs, which puts a free speech layer on the same workday Flow wants to own. If Rambler is good enough for a Slack ping, Wispr has to be good enough for the document you cannot afford to send wrong, and it has to do that in the room you are actually in.

An Alexa Veteran Is Building Past Dictation

The personnel choice is the tell. On July 24, 2026, Wispr opened the Advanced Interfaces Lab under chief scientist Ariya Rastrow, a founding member of the Amazon team behind Alexa who later led voice and multimodal work at Meta. The lab page says the group builds the interface between people and intelligence, with four jobs on the wall: contextual speech models with long memory, voice in and UI out, routing a request to the right model, and help that shows up before you ask.

Rastrow’s line is blunt about what the lab is not building.

The goal isn’t for AI to talk back to you. It’s to understand intent and generate the right interface for the job.

Ariya Rastrow, Chief Scientist, Wispr Flow

Kraning’s reading of that hire is that the last generation of voice failed on intelligence, and that now the intelligence exists, the old speech-to-text-to-model pipe is the wrong shape. The lab’s stated swap is voice-to-outcome: an action, a workflow, or a screen built for the moment, rather than a chatbot paragraph. Engineers already talk to coding agents because typing prompts at a system that types faster than they do is a bottleneck. Sales reps, in Menlo’s sketch, talk through a call and the CRM updates itself.

The first product past the cursor is already out. Wispr Flow Notetaker launched on August 5, 2026, on Mac, with Windows planned, as a meeting recorder that names speakers and turns talk into notes you can use. Kothari called Flow what you say to your computer, and Notetaker what you say to everyone else. Both, he said, get better as the models under them do.

Why Hinglish Is Canto’s Hardest Test

The other half of Canto is not noise. It is how people actually talk. Kothari wrote that roughly half the world moves between languages during a normal day, often inside a single sentence, and that models trained on one language at a time handle that badly. Sahaj Garg, Wispr’s chief technology officer and co-founder, put the same split at about 50% of people in a January 2026 research note, and said Flow already routes some languages to different speech engines because standard models fail on Hindi, Marathi, Thai, and Tamil.

Hinglish is the worked example. Hindi speakers usually write in Devanagari, but when they speak Hinglish they want romanized script they can send. Hearing every word is not enough if नहीं comes back in the wrong alphabet. Flow’s help docs say Hinglish output is always Latin script, with community spellings such as “nahi” and “arey,” dandas turned into periods, and a warning that English loanwords outside a common-word list can come back as phonetic misses like “sttej” for stage. Users are told to pick Hinglish on purpose, not Hindi or English alone, and to keep the language list to two or three for accuracy.

That is a product requirement, not a localization sticker. New checks in the Series B came from Peak XV, Together Fund, and Activate, funds that know those mixed-language days as a market, not a demo. Domantas Sabonis, the three-time NBA All-Star who joined the round, said he uses Flow every day and grew up rotating among English, Spanish, and Lithuanian. “It keeps up with me no matter which one I’m speaking,” he said. “So the opportunity to invest and support a product I use and believe in was a no brainer.”

The Cap Table Jumped to $2 Billion

Kothari and Garg founded Wispr in 2021 after meeting as Stanford batchmates. Kothari has said the long bet started in Delhi, after he watched Iron Man as a child and wanted a computer that understood him. The company is now selling that bet at almost three times the price secondary listings put on it last autumn.

HOW THE ROUND STACKED UP

  1. 2021: Tanay Kothari and Sahaj Garg found Wispr in San Francisco.
  2. June 24, 2025: Menlo leads a $30 million Series A.
  3. November 20, 2025: Notable Capital leads a $25 million extension that takes total funding to $81 million, listed at $700 million post-money; Hans Tung of Notable joins as a board observer.
  4. July 24, 2026: Rastrow’s Advanced Interfaces Lab opens.
  5. August 5, 2026: Notetaker ships on Mac.
  6. August 17, 2026: Menlo leads the $280 million Series B at a $2 billion valuation; total capital reaches $361 million.

Existing backers Notable Capital, NEA, Neo Ventures, 8VC, and MVP Ventures returned. New institutional names include Acrew, Activate, Forerunner, Goodwater, Peak XV, Together Fund, and PLUS Capital, plus a Plus Capital group of athletes and cultural figures that includes Shaun White, Joe Burrow, Klay Thompson, Paul George, and Sabonis. Kraning called the check one of the largest investments Menlo has made in an AI company, and said the first keyboard bet, placed about 14 months earlier, had paid off.

WHERE EXPERTS DISAGREE

  • Matt Kraning, Menlo: The bottleneck in AI has moved from the model to the interface, and Wispr is that interface; “free and mostly right” is not enough for people who talk for a living.
  • Michael Ashley Schulman, Cerity Partners: The value may be there if the growth rate survives, because triple-digit quarterly growth off a small base is easy for a few quarters and hard once early adopters stop being the whole customer base.

The multiple from $700 million to $2 billion in nine months is the part that still has to show up in untouched sentences, not in the press note. Canto is a preview. Rambler is already on a flagship phone. Flow has to make the next word come back the way it was meant, in the room where it was said, or that $2 billion is just a very expensive microphone.

Disclaimer: This article is news reporting and analysis of a private company’s funding round and product plans, and it is for information only. It is not investment advice, a solicitation to buy or sell securities, or a recommendation of any venture, secondary sale, or fund. Readers who are considering an investment in private tech companies should consult a licensed financial adviser or securities lawyer who can review their own facts. Figures for valuation, revenue growth, error rates, and product status come from company posts, investor notes, and listings as of the dates named above and can change as later rounds, model releases, or filings appear.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending