Private AI: from the cautious option to the default
Open models have come close to the best, and organizations worldwide are bringing sensitive AI home. In Iran, keeping work running is at stake too.

Until a few years ago, running AI on your own servers meant settling for weaker models. In 2026 that trade-off has changed. Models you can run on your own hardware are close to the best in the world, and large organizations everywhere are moving sensitive AI work back onto infrastructure they control. For organizations in Iran there is a third reason, more urgent than the other two: keeping work running.
Open models have closed most of the gap
An open model is one whose maker lets others download it and run it on their own servers. Stanford's AI Index 2026 shows that on the Chatbot Arena leaderboard, where users compare answers without knowing which model wrote them, the best closed model's lead over the best open one fell from 15.2% in May 2023 to about 3% in March 2026. Epoch AI estimated in May 2026 that the best open models trail the best closed ones by about four months on average.
Cost fell just as fast. According to Stanford's 2025 report, the cost of running a model at GPT-3.5 level dropped more than 280-fold between November 2022 and October 2024. Epoch AI found that models running on a single consumer graphics card (under USD 2,500) match what the best models could do 6 to 12 months earlier. Many of them, including gpt-oss, Qwen3, DeepSeek, Mistral 3 and Gemma 4, are released under permissive licences such as Apache 2.0 or MIT that allow running them on your own servers. Read each licence before use: some of the newest very large models come with terms of their own.
Organizations worldwide are bringing sensitive data home
Gartner named geopatriation among its top strategic trends for 2026: moving data and applications out of global public clouds into sovereign clouds, regional providers or the organization's own data centre. Gartner expects more than 75% of enterprises in Europe and the Middle East to have done so by 2030, up from less than 5% in 2025.
- In Deloitte's State of AI in the Enterprise 2026, which surveyed 3,235 leaders in 24 countries, 83% said sovereign AI is at least moderately important to their strategic planning, and 77% factor the country of origin of an AI solution into vendor selection.
- In IBM's Cost of a Data Breach Report 2026 (July 2026), the share of organizations with security incidents involving shadow AI, meaning tools employees use without the organization's approval, rose from 20% to 43% in a year, and 68% of breached organizations had no AI governance policy.
- In August 2026 Gartner predicted that the inference cost of each agentic workflow will rise more than fivefold through 2028, even as the price per token falls, because agents call models many times per task. Our reading: fixed capacity on your own hardware makes that cost more predictable.
In Iran, continuity is at stake
Iran is not on the supported-country lists of OpenAI, Anthropic or Google's Gemini (checked on 29 September 2026). Any workflow that depends on these services can stop without notice, or have its account suspended.
International connectivity is not guaranteed either. Cloudflare Radar data shows Iran's internet traffic stayed below 1% of normal for about three months from 28 February 2026, and only part of it came back afterwards. Several multi-day shutdowns were also recorded in June 2025. During those periods, every organization whose AI ran outside the country lost it.
Domestic rules point the same way. The Ministry of Science's guideline on AI in research (September 2026) classes data as public, sensitive, confidential or secret: confidential data goes into public AI tools only with explicit permission, and secret data never. The National AI Development Law, issued in September 2026, stresses support for open-source technology and preparedness for possible sanctions.
What private AI means, and what it takes
Private AI is not just a model installed on a server. In our experience, five parts have to work together:
| Part | The question to answer |
|---|---|
| Hardware | How many users at once, with which models, at what response time? The server and graphics cards follow from that. |
| Model | Which open model suits Persian and your work, and does its licence allow your use? |
| Connection to your knowledge | Which documents does the assistant answer from, and does it show the source of each answer? |
| Access and control | Which documents and which model can each user reach? |
| Operations | Who updates the models, monitors the system and supports users? |
An organization that answers these five questions before buying ends up with an everyday tool, not a pilot that gets shelved.
Vakav's view
We built Vakav Autonomous on the premise that an organization's data should not leave it, and its work should not depend on an outside connection. Autonomous runs entirely on the organization's own servers with no internet, gives each department its own workspace with access that follows the org chart, answers from the organization's own documents with sources cited, and lets you choose a different model for each workspace. Bots and custom flows automate repetitive work.
To see what private AI would look like in your organization, visit the Vakav Autonomous page or start a pilot with us.