HomeCompanyPortfolioServicesSoftware & AIMobile AppsIndustriesLocationsPricingBlogContact
Englishالعربية
Home  /  Blog  /  AI
AIJan 17, 2026·10 min read

Generative AI in the Design Studio: Where It Helps and Where It Hurts

IW
IITWares Editorial Team
Digital Strategy & Search
Generative AI in the Design Studio: Where It Helps and Where It Hurts

A working guide to generative ai design for companies operating in Saudi Arabia — grounded in local search behaviour, local regulation and what we see across client accounts.

The reliable returns from AI right now are unglamorous — support deflection, document search, drafting, extraction. The ambitious autonomous systems work best once those foundations exist and the data underneath them is clean.

What good looks like here

The competitive picture matters more than the checklist. Before committing to generative ai design, look at who is currently visible for your commercial terms, how strong they actually are, and whether the results page is dominated by aggregators. In several Saudi B2B and industrial categories the first page is still thin, and a well-executed programme reaches it within a quarter. In retail, real estate and travel, expect a considerably longer campaign.

Governance, in one page

Which tools are approved. What data may never be pasted into an external model. When AI assistance must be disclosed. Who reviews AI output before it reaches a customer. How incidents are reported. One page people actually read beats a policy document that lives unopened on the intranet — and given SDAIA's active role in AI and data governance, having something written is now table stakes.

Choose the cheapest architecture that solves the problem

Prompt engineering with a capable general model handles more than most teams expect. Retrieval-augmented generation adds your own documents and is the right answer for the majority of business use cases. Fine-tuning is for consistent format, tone or a narrow specialised task — rarely for adding knowledge. Work upward through that ladder and stop at the first rung that meets the requirement; each step up multiplies cost and maintenance.

Arabic changes the engineering

Arabic performance varies considerably more between models than English performance does, dialect handling is uneven, and tokenisation is less efficient — meaning higher cost per equivalent output. Retrieval quality also suffers if your embedding model handles Arabic poorly. Evaluate on your own Arabic content with your own questions before committing; published English benchmarks will mislead you here.

The businesses that win in Saudi search are rarely the biggest. They are the ones that did the unglamorous work consistently for four quarters.

Evaluation before deployment

Build a test set of a hundred real questions with known good answers before launch. Score accuracy, refusal behaviour on out-of-scope questions, and tone. Re-run it whenever you change the prompt, the model or the corpus. Without this you are shipping on anecdote, and quality regressions arrive silently after routine changes.

Retrieval quality is the whole system

Most disappointing AI deployments are retrieval failures wearing a generation costume. If the right passage is not fetched, no model can answer well. Invest in document preparation, sensible chunking, metadata, hybrid keyword-plus-vector search and re-ranking. Measure retrieval separately from generation so you know which half is failing.

Corroboration beats assertion

Generative systems weight claims that appear consistently across independent sources. A price stated only on your own website is an assertion; the same price reflected in a directory listing, a press mention and a third-party review becomes a fact. Invest in being described accurately elsewhere — trade media, chambers, industry associations, partner sites — because that off-site consistency is what converts your content into citable material.

Typical pilot shape

StageTypical windowWhat you should see
Use case selection and baseline1–2 weeksMust be measurable or the pilot cannot be judged
Data preparation and retrieval build2–4 weeksUsually the largest share of effort
Evaluation and tuning2–3 weeksAgainst a hundred-question test set
Controlled production rollout4–8 weeksWith human review on defined risk thresholds

Windows assume consistent execution and a market of ordinary competitiveness. Treat them as planning ranges, not commitments.

Generative engine optimisation, defined without hype

GEO is the practice of making your content the material a generative system reaches for when composing an answer. It shares its foundations with SEO — crawlability, authority, clarity — but shifts the objective from position to inclusion. Success looks like being named in a synthesised paragraph rather than sitting at position three. The tactics are less exotic than the label suggests: be retrievable, be quotable, be corroborated.

Human in the loop, positioned deliberately

Decide in advance which decisions the system may take alone, which need approval, and which it must never take. Log every action for audit. Set confidence thresholds that escalate rather than guess. This is what makes automation defensible to auditors, regulators and the team whose work it touches — and it is what keeps a small error from becoming a systemic one.

A checklist you can run this week

The national context, briefly

Saudi Arabia designated 2026 its Year of Artificial Intelligence, with substantial state-backed investment channelled through SDAIA, sovereign AI vehicles including HUMAIN, and Arabic-language model development such as ALLaM. For an ordinary business the significance is less about the headline figures and more about what they produce downstream: local compute capacity, in-Kingdom cloud regions, Arabic models that work properly, a talent pipeline, and procurement expectations that increasingly assume digital maturity.

Community and third-party surfaces

Forums, Q&A threads, review platforms and community discussions are disproportionately represented in AI answers because they contain candid, experience-based language. Participating honestly — answering questions in your field under a real identity, without spamming links — puts your expertise into exactly the sources these systems favour. This is slow, human work and it is difficult for a competitor to copy quickly.

Where to start this week

Choose one contained use case with a measurable baseline — support deflection, document search, invoice extraction. Build a hundred-question evaluation set from real examples before you build anything else. Test your shortlisted models on your own Arabic content rather than published benchmarks. Write the one-page usage policy while the pilot runs.

Pick the two changes above with the clearest link to revenue and ship them this month. Momentum matters more than completeness at the start, and a finished small change beats a planned large one.

[ Key Takeaways ]
Typical pilot shape
Arabic changes the engineering
Choose the cheapest architecture that solves the problem
Evaluation before deployment
Share

Frequently asked questions

How do we handle Arabic properly?+

Evaluate candidate models on your own Arabic content with your own questions. Arabic performance varies far more between models than English performance, and tokenisation makes it more expensive per equivalent output.

Can we keep data inside the Kingdom?+

Yes — through local hyperscaler regions with contractual guarantees, sovereign cloud arrangements, or self-hosted open-weight models. The trade-off is capability and cost against control.

RAG or fine-tuning?+

RAG for adding your own knowledge, which covers most business cases. Fine-tuning for consistent format, tone or a narrow specialised task. Start with prompting and move up only when it demonstrably fails.

What does an AI pilot cost?+

A contained, well-scoped pilot with a clear baseline is usually a five-figure riyal investment over six to eight weeks. Costs escalate when scope is vague and no baseline exists to judge success against.

Can you work in Arabic and English?+

Yes — both languages natively, across strategy, content, design and development, which is generally where translated-only providers run into trouble.

Keep Reading

All Articles →

Ready to put these ideas to work?

Start a Project →