Cantonese AI: Hong Kong's Votee Takes on English-Mandarin Dominance
Source: Fortune. Casualplayhub News adds summary, context, and editorial framing while linking back to the original report.
The artificial intelligence revolution is overwhelmingly a two-language story. English and Mandarin Chinese power the most advanced models from nearly every major developer—OpenAI, Anthropic, DeepSeek, Moonshot AI, and Z.ai among them. These companies, all based in the United States or mainland China, have built systems that excel in their native tongues. But that leaves a vast linguistic landscape behind, including languages spoken by tens of millions of people.
"The whole AI revolution is in English and Mandarin," says Pak-Sun Ting, CEO of Hong Kong-based startup Votee AI. "There's only a very small fraction that represents other languages." His company is determined to change that. Votee takes open-weight models from developers like Meta and Alibaba, retrains them on Cantonese data, and sells the resulting AI to banks, universities, and government departments.
Cantonese is often dismissed as a mere dialect of Chinese, but it is profoundly different from Mandarin. It uses distinct grammar, vocabulary, and in Hong Kong, speakers frequently mix English and Cantonese within the same sentence. More than 80 million people speak Cantonese—roughly the same number as Korean speakers, and more than Italian or Thai. Yet the language lacks a deep pool of standardized written data, especially for colloquial usage.
Leading models can handle basic Cantonese, but they routinely stumble on cultural and local knowledge, according to HKCanto-Eval, a benchmark set developed by researchers at Kyushu University, the Education University of Hong Kong, and the local AI community hon9kon9ize, with sponsorship from Votee. Ting explains that building a Cantonese LLM is "essentially taking the same steps as if you were training a model from scratch." The process involves starting with an existing open-source model like Meta's Llama or Alibaba's Qwen, then performing additional training with Cantonese data.
Votee gathers its Cantonese data through online scraping, including content from Radio Television Hong Kong (RTHK), the city's public broadcaster. It also receives data from universities and the community, and draws on its own background as a big data company. To supplement, the startup uses synthetic data, creating its own Cantonese datasets. These efforts have expanded the corpus from 100 million tokens to over 500 million.
Votee's models are around 70 billion parameters—significantly smaller than the best frontier models. Yet Ting insists they are capable enough to understand and reason in Cantonese. More importantly, the training costs are far lower. Ting estimates the company uses between 500 million and 1 billion tokens for training, compared to the trillions used for English-language models, at a cost of roughly $250,000.
Votee is not alone in this mission. Several companies are building models for so-called "low resource languages." Indonesia's Indosat is developing Sahabat AI for Bahasa and other Indonesian languages. Singapore's state-backed AI Singapore runs SEA-LION, covering 11 under-resourced Southeast Asian languages. South Korea has gone further, staging a state-sponsored elimination tournament dubbed "AI Squid Game" to pick national champions, backed by a 2026 AI budget of about $6.8 billion.
These efforts fall under the banner of "sovereign AI"—the idea that governments and companies should own their own data, models, and infrastructure instead of renting them from abroad. "AI has become such an essential need, and so you don't want to be tethered to anybody else who can turn it off," Ting says. He acknowledges that full sovereign AI—owning every part of the supply chain—is "very difficult." Instead, he suggests countries focus on owning foundation models and the applications built on them. Governments often do not need the most powerful frontier model; a tiny model with as few as 1 billion parameters can automate routine tasks. For advanced needs, they can route a powerful English or Chinese model's outputs through a smaller local-language layer.
Ting notes that Votee works with models from MiniMax and SenseTime, and can use Nvidia chips. "We can use Nvidia chips, we can use Moonshot or DeepSeek's model," he says. "We're that person in high school who's friends with everyone."
Despite his talk of preservation, Votee is a for-profit company. Governments and corporations are the first customers for AI in languages like Cantonese or Bahasa. Ting says the startup is profitable "in the sense that our revenues exceed our costs," funded largely through client contracts. Votee now counts Hong Kong tycoon Allan Zeman, who developed the Lan Kwai Fong nightlife district, as an advisor.
Votee's ambitions extend beyond Hong Kong. Ting says the startup is in "active discussions" with AI Singapore and plans to expand further into Southeast Asia. Beyond that, he wants to explore using AI to protect endangered languages in East Asia, North America, and Africa.
Ting calls the dominance of English-language AI a "typewriter moment"—a productivity gain so large that people abandon their own language to get it. "People will adopt English just because the typewriter's productivity is so strong versus their own language," he says. Whether a 70-billion-parameter Cantonese model can reverse that trend remains uncertain. But Ting is driven by a desire to give other languages a fighting chance. "Every language that dies, you lose another way of seeing the world," he says. "That could just be preserved in a museum where you can kind of see it. But we can also unlock a lot of new wisdom."
Article commentary
The rise of English and Mandarin as the twin engines of AI development is a natural consequence of market forces, but it also creates a significant cultural and economic blind spot. Votee AI's effort to build a Cantonese model highlights a critical tension: the world's most powerful technology is becoming increasingly inaccessible to hundreds of millions of people who do not speak either of the dominant languages. Cantonese, with over 80 million speakers, is a prime example of a language that is numerically large yet linguistically marginalized in the AI ecosystem. What makes Votee's approach noteworthy is its pragmatism. Rather than attempting to build a foundational model from scratch—a capital-intensive endeavor that would rival the budgets of OpenAI or DeepSeek—the company repurposes existing open-weight models and fine-tunes them with curated local data. This strategy dramatically lowers costs and time to market, making it replicable for other underserved languages. The 70-billion-parameter model is dwarfed by frontier systems, but as Ting rightly points out, many government and enterprise applications do not require the raw power of a trillion-parameter model. A focused, smaller model tailored to local needs can be more effective than a generic giant. The concept of sovereign AI is gaining traction globally, and Votee's work fits neatly into that narrative. Governments are increasingly wary of relying on foreign AI infrastructure, especially for sensitive domains like public administration, healthcare, and defense. The ability to run a local model—even a less capable one—on local hardware offers a degree of autonomy and security that renting from overseas cannot provide. However, the full sovereign AI vision is fraught with challenges. Most countries lack the computational resources, talent, and data to build and maintain frontier-level models. Ting's suggestion to focus on owning the application layer while routing through powerful foreign models for heavy lifting is a realistic middle ground. Culturally, the stakes are high. Language is not just a communication tool; it carries unique worldviews, idioms, and knowledge systems. As AI becomes ubiquitous, the risk of language shift—where speakers abandon their native tongue for the convenience of AI—is real. Ting's "typewriter moment" analogy is apt: historical technological shifts often accelerated the dominance of English, from the printing press to the internet. AI could be the next such force. By investing in local language models, companies like Votee are not only building a business but also preserving linguistic diversity. Yet there are limits. Votee's model relies heavily on data from public broadcasters, universities, and synthetic generation. For languages with even smaller speaker bases or less digitized content, this approach may not scale. The economics also remain uncertain: while Votee claims profitability, its customer base is currently limited to government and corporate clients. Expanding to consumer markets for endangered languages may require subsidies or philanthropic support. Moreover, the quality of a 70-billion-parameter model may not satisfy all users, especially as frontier models continue to improve rapidly. Overall, Votee AI represents a promising but incomplete solution. It demonstrates that local language AI is feasible and valuable, but it also underscores the asymmetries in the global AI landscape. The real test will be whether such models can achieve widespread adoption and whether they can keep pace with the relentless advancement of English and Mandarin-centric systems. In the meantime, efforts like these are essential for ensuring that the AI revolution does not leave entire linguistic communities behind.