essay · 10 August 2026 · 9 min read
The Bottleneck Is Not Compute
Saudi Arabia is buying AI infrastructure at historic scale, and a frustrated Saudi user just explained better than any analyst why the chips are the easy part: the Arabic training corpus is thin, half of it is translated English, and the model's own rules forbid it the conversations its users actually want.
Last week a Saudi user named @AboDantee posted a long thread about HUMAIN Chat, the flagship chatbot of Saudi Arabia’s state AI venture, after using it for seven days. He was not gentle. The product, he wrote, is a failure, and no Arabic model will succeed “with this quantity of restrictions and biased information,” built on Arabic content that is “weak, flimsy, and without objectivity.” The image he reached for has stayed with me: all the GPUs in the world cannot help “if the information is corrupted — it is as if you had a powerful computer with an RTX 5090 and endless RAM, and in the end you installed DOS on it.” (The translations here are mine.)
It would be easy to file this under ordinary product complaints, and some of it is that: he also lists login failures and a backend that “isn’t working.” But the thread is something better. It is the clearest statement I have read of why the sovereign-AI wave now cresting in the Gulf has its economics backwards, written not by a think-tank skeptic but by exactly the customer the project was built for. I wanted to argue with him. I work on AI systems for this region myself, and I would like the sovereign-AI story to end better than a week of goodwill and a shrug. He is right anyway.
The bet he is criticizing is enormous. HUMAIN launched in May 2025 as a company of the Public Investment Fund, chaired by the crown prince, with a mandate to build the full AI stack inside the Kingdom. Within days it had announced a partnership with Nvidia for AI factories of up to 500 megawatts over five years, beginning with an 18,000-GPU Grace Blackwell supercomputer; a collaboration with AMD valued at up to ten billion dollars; and a five-billion-dollar-plus “AI Zone” with Amazon Web Services. By late 2025 the company was describing its existing data-center capacity as sold out and targeting roughly six gigawatts by 2034; by November it had added a 500-megawatt project with Nvidia and xAI. Whatever else is said below, none of it is a claim that this infrastructure is fake or useless. The machine is real.
The question is what the machine has to eat. HUMAIN Chat, launched in August 2025 as an Arabic-first assistant for what the company frames as four hundred million Arabic speakers, runs on ALLaM, an Arabic model family developed with the Kingdom’s national data authority. And the most telling document about it is not a review but the ALLaM paper itself, which reports, with admirable candor, that the team curated 540 billion Arabic tokens for pretraining, “of which 270B are natural Arabic tokens and 270B are translated Arabic tokens.” Half of the sovereign Arabic model’s Arabic is machine-translated English. The paper is equally direct about why: pretraining data in Arabic “is much more limited” than in English, where some four trillion tokens were available to the same team.
This is not a Saudi peculiarity; it is the condition of the language online. Arabic is used daily by more than 400 million people, roughly five percent of humanity, and accounts for 0.6 percent of websites; in the latest Common Crawl, the open web scrape underlying most large models, Arabic pages are outnumbered by English ones about sixty to one. The UAE’s Jais team, building the region’s other flagship Arabic model, wrote it out even more bluntly: after collecting everything they could, they had 72 billion unique Arabic tokens and concluded that by standard scaling laws “we do not have enough data to train a 13B parameter Arabic model” — so they, too, padded the corpus with translation. When Abu Dhabi’s TII launched Falcon Arabic in 2025, its headline selling point was that the training data was native and non-translated, which tells you what everyone in the field understands the default to be. @AboDantee’s engine-and-tires image is not a metaphor the industry would dispute. It is a summary of the industry’s own papers.
Thinness is only the first problem, and his thread is careful to separate it from the second. He asked the model religious questions and was refused; tribal questions, refused or answered with what he calls laughably wrong information. A safety system, he writes, should stop what is dangerous, “not hide part of history, or political topics, or religion, or logical analysis.” No independent audit of HUMAIN Chat’s refusal behavior has been published yet, so his account is testimony rather than measurement, and it should be read that way. But the design intentions are documented where you would expect them. HUMAIN’s own terms of use prohibit content that violates “public morals, Islamic values, or national security interests of the Kingdom of Saudi Arabia” — the launch coverage led with “Islamic values” in the headline — and the Jais team stated openly that their chat model was trained to “avoid engaging in discussions on sensitive topics, particularly those related to certain aspects of religion and politics.” Nor is the pattern regional. A study in PNAS Nexus this February found Chinese-built models refusing politically sensitive Chinese-language questions at rates between roughly a third and 60 percent, against near-zero for their Western counterparts, and the Oversight Board’s July assessment of ten models found refusal rates of 34 percent for political criticism touching restrictive jurisdictions — Saudi Arabia among them — against 14 percent elsewhere. National models inherit national red lines; that is close to a law of the genre.
What @AboDantee adds, and what the ministries commissioning these systems seem not to have priced in, is how this reads from the user’s chair. A model that backs away from every sensitive question, he writes, will be seen as a sign of stupidity “even if you have a million Nvidia chips — in the end this is a stupid model.” He is describing something real about how capability is perceived: a refusal is indistinguishable, from the outside, from ignorance. And his most cutting observation follows directly from it. ChatGPT and Gemini, he says, understand Arabic and Saudi culture better than the Saudi model does, and answer “with full credibility.” There is no published head-to-head to confirm the comparison, but the mechanism that would produce it is exactly the one this essay has been describing. The foreign models train on everything the Arabic web could not hold: the English-language scholarship on the region, the diaspora press, the arguments that happen off the record of the domestic corpus — and they operate under fewer of the red lines that a domestic owner must impose on itself. A sovereignty implemented as restriction can end with the national model knowing its own country less well than the imports do.
Step back from HUMAIN for a moment, because his complaint points at something the whole field keeps mispricing. The metric that finally decides a product is the experience of using it, the entire encounter and not the paint job: whether it answers the question you actually asked, whether it earns a second visit, whether it leaves you feeling capable or stupid. For an AI model that experience is not a layer on top of the product; it is the product. And the industry has spent three years optimizing the things that photograph well on a slide, the parameter counts and benchmark scores and gigawatts, because those are the numbers you can measure and put in a press release. @AboDantee is measuring the other thing, the one that never makes the announcement: he opened the app every day for a week and it did not help him. His own image for it is a Lamborghini engine bolted onto bicycle tires, all the horsepower in the world geared to a ride that shakes itself apart. A model can top every Arabic leaderboard and still lose, decisively, in the only benchmark that compounds, which is tomorrow morning, when the user has a real question and has to choose which icon to open.
Chinese models carry heavier documented censorship than anything alleged of HUMAIN, and it has not stopped their global adoption: of the fifty most-used generative-AI apps, twenty-two are Chinese, most of them used mainly abroad. But that is precisely the point. DeepSeek wins users in languages and on topics where its red lines rarely bind. An Arabic-first assistant for Saudis lives its entire life exactly where its red lines bind — religion, politics, history, the tribal genealogies its users actually ask about. The Chinese playbook of exporting around your own restrictions is not available to a product whose whole market sits inside them.
The deeper trouble with compute-first sovereignty is that the scarce input was never for sale. Nvidia’s Jensen Huang, who did more than anyone to popularize the phrase, told a Dubai audience that sovereign AI “codifies your culture, your society’s intelligence, your common sense, your history”. But a model can only codify the culture that made it into machine-readable text, and the analysts who have looked closely — Carnegie calls full sovereignty “beyond the reach of even the most capable middle power actors,” Brookings calls the full stack “structurally infeasible,” AI Now calls the compute route “a fantasy of independence” — are converging on the same conclusion from the supply side that @AboDantee reached from a chat window. Chips are importable. An information environment is not. It has to be grown: digitized archives, a plural press, scholarship in the language, criticism that survives long enough to be crawled. Egypt, Sweden, and Silicon Valley all feed the same crawlers; what differs is how much each society has allowed onto the record.
None of this makes the Saudi buildout foolish, and the thread’s author is not arguing that it is. His constructive demand is the sensible one: point the gigawatts at industrial, medical, and government systems where compute genuinely is the bottleneck, and stop spending them on “a meaningless chat.” If the conversational ambition is serious, the work that would change the outcome is slower and less photogenic than a GPU purchase — funding native-Arabic corpus construction the way Falcon Arabic has begun to, digitizing what the region’s libraries and newspapers hold, disclosing plainly what user conversations are used for, and letting the corpus keep enough of its arguments to be worth learning from. He closes with a warning any product manager would recognize: patriotism buys a week of goodwill, maybe a month, “and then I will go back to using the product that actually helps me.” The chips arrived in eighteen months. The well they need takes a generation, and no one has yet found a way to import one.
Comments