Hong Kong startup Votee AI trains Cantonese models to counter English-Mandarin AI bias

Votee AI, a Hong Kong startup, retrains open-weight models from Meta and Alibaba on Cantonese data to serve banks, universities, and government departments. CEO Pak-Sun Ting argues that AI is useless without coverage of local languages like Cantonese, which has over 80 million speakers but lacks standardized written data. The company is part of a growing effort to address the neglect of low-resource languages in the AI boom.
Votee AI’s approach relies on retraining existing open-weight models rather than building from scratch, using a mix of scraped public broadcasts, university contributions, and synthetic data. This expanded their Cantonese corpus from 100 million to over 500 million tokens, with training costs around $250,000—far below frontier labs. Their models, roughly 70 billion parameters, serve local institutions needing accurate handling of Cantonese in education, healthcare, and policing.
The company’s work parallels broader regional efforts, such as Indonesia’s Sahabat AI and Singapore’s SEA-LION, all targeting under-resourced languages. Hong Kong’s unique code-switching between English and Cantonese adds complexity, while benchmarks like HKCanto-Eval show mainstream models still stumble on local cultural knowledge, underscoring the niche Votee fills.
This story could reshape how non-English, non-Mandarin speakers access AI services. If successful, Votee’s models may enable more accurate automated tools for Cantonese-speaking communities in banking, public services, and education, reducing reliance on flawed translations. However, the impact may remain limited to institutional clients initially, and broader societal adoption could hinge on cost and accessibility. The effort also highlights a growing divide: regions without such initiatives risk being left behind in the AI-driven economy.