HomeCompanyPortfolioServicesSoftware & AIMobile AppsIndustriesLocationsPricingBlogContact
Englishالعربية
Home  /  Blog  /  AI
AIApr 04, 2026·12 min read

Fine-Tuning vs Prompting vs RAG: Choosing the Cheapest Thing That Works

IW
IITWares Editorial Team
Digital Strategy & Search
Fine-Tuning vs Prompting vs RAG: Choosing the Cheapest Thing That Works

This is a field guide to fine tuning vs rag for the Saudi market. No theory you can't act on, and no advice that assumes a US search landscape.

Saudi Arabia declared 2026 its Year of Artificial Intelligence, and the investment behind that is real. What matters for an individual business is narrower: which capabilities can be bought and operated today, in Arabic, at a cost that pays back.

Framing the problem properly

Most teams arrive at fine tuning vs rag after something stopped working: enquiries fell, a competitor became visible, or a target was missed. That context matters, because the right first move differs depending on whether you are fixing a decline or building from a standing start. Diagnose which situation you are in before applying anything below — the sequence changes completely, and applying a growth playbook to a decline problem wastes a quarter.

Governance, in one page

Which tools are approved. What data may never be pasted into an external model. When AI assistance must be disclosed. Who reviews AI output before it reaches a customer. How incidents are reported. One page people actually read beats a policy document that lives unopened on the intranet — and given SDAIA's active role in AI and data governance, having something written is now table stakes.

Where the return actually shows up

The reliable wins are unglamorous: first-line support deflection, document search across years of accumulated files, drafting and summarising routine correspondence, extracting structured data from invoices and forms, and translation quality assurance. Each is measurable, contained and pays back inside a year. The ambitious autonomous agent projects usually work best after these foundations exist.

Cost control from day one

Token costs scale with usage in ways that surprise finance teams in month three. Cache repeated queries, route simple requests to smaller models, cap context length, monitor per-feature spend, and set alerts. Design cost observability in at the start; retrofitting it once a system is embedded in daily operations is considerably harder.

Choose the cheapest architecture that solves the problem

Prompt engineering with a capable general model handles more than most teams expect. Retrieval-augmented generation adds your own documents and is the right answer for the majority of business use cases. Fine-tuning is for consistent format, tone or a narrow specialised task — rarely for adding knowledge. Work upward through that ladder and stop at the first rung that meets the requirement; each step up multiplies cost and maintenance.

Visibility is no longer a position on a page. It is whether the machine composing the answer considers you a source worth naming.

Evaluation before deployment

Build a test set of a hundred real questions with known good answers before launch. Score accuracy, refusal behaviour on out-of-scope questions, and tone. Re-run it whenever you change the prompt, the model or the corpus. Without this you are shipping on anecdote, and quality regressions arrive silently after routine changes.

Managing AI crawlers deliberately

GPTBot, ClaudeBot, PerplexityBot, Google-Extended and others can each be allowed or blocked in robots.txt. Blocking protects content from training use; it also removes you from the answers those systems produce. For most Saudi service businesses seeking visibility, allowing access to public marketing pages while excluding client portals, gated assets and internal search results is the sensible middle position. Decide it consciously rather than inheriting a default.

Typical pilot shape

StageTypical windowWhat you should see
Use case selection and baseline1–2 weeksMust be measurable or the pilot cannot be judged
Data preparation and retrieval build2–4 weeksUsually the largest share of effort
Evaluation and tuning2–3 weeksAgainst a hundred-question test set
Controlled production rollout4–8 weeksWith human review on defined risk thresholds

Windows assume consistent execution and a market of ordinary competitiveness. Treat them as planning ranges, not commitments.

What actually changes for a mid-market company

Three practical effects. Local infrastructure lowers latency and simplifies data residency arguments. Better Arabic models make customer-facing automation viable where it previously was not. And rising expectations mean clients and government buyers increasingly assume you can transact digitally. That last one is the competitive pressure most companies feel first.

Baseline before pilot, always

Record current cycle time, error rate, cost per transaction and volume before you deploy anything. Without that baseline the review meeting becomes a debate about impressions. With it, the conversation is arithmetic — and arithmetic is what unlocks funding for the next phase.

Retrievability: can a machine actually read you?

Many AI crawlers do not execute JavaScript, do not wait for lazy-loaded content and do not scroll. If your key facts live inside a tab, an accordion opened by script, an image, or a client-rendered component, they may as well not exist. Put the substance in server-rendered HTML. Provide text alternatives for anything visual. Test by fetching your page as raw HTML and reading what comes back.

The working checklist

The talent picture

Demand for data engineers, ML practitioners, cloud architects and AI-literate product people substantially exceeds local supply, which raises salaries and lengthens hiring cycles. Saudization targets add a further constraint. The pragmatic responses are training existing staff, partnering with a specialist provider for the build while developing internal capability to operate it, and designing systems that do not require rare expertise for routine maintenance.

Entities, not just keywords

Modern systems reason about things: your company, your founders, your services, your locations, your clients. Strengthen those entities with consistent naming, sameAs links to every official profile, Organization schema, a substantive About page with founding date and leadership, and Wikidata or industry-database presence where you legitimately qualify. A well-defined entity gets recommended; an ambiguous one gets skipped.

Where to start this week

Choose one contained use case with a measurable baseline — support deflection, document search, invoice extraction. Build a hundred-question evaluation set from real examples before you build anything else. Test your shortlisted models on your own Arabic content rather than published benchmarks. Write the one-page usage policy while the pilot runs.

The Saudi market is moving quickly enough that a decision deferred by two quarters is usually a decision made by a competitor instead. Choose the smallest useful version and start.

[ Key Takeaways ]
Entities, not just keywords
The talent picture
Framing the problem properly
What actually changes for a mid-market company
Share

Frequently asked questions

Can we keep data inside the Kingdom?+

Yes — through local hyperscaler regions with contractual guarantees, sovereign cloud arrangements, or self-hosted open-weight models. The trade-off is capability and cost against control.

How do we handle Arabic properly?+

Evaluate candidate models on your own Arabic content with your own questions. Arabic performance varies far more between models than English performance, and tokenisation makes it more expensive per equivalent output.

RAG or fine-tuning?+

RAG for adding your own knowledge, which covers most business cases. Fine-tuning for consistent format, tone or a narrow specialised task. Start with prompting and move up only when it demonstrably fails.

What does an AI pilot cost?+

A contained, well-scoped pilot with a clear baseline is usually a five-figure riyal investment over six to eight weeks. Costs escalate when scope is vague and no baseline exists to judge success against.

How much does this cost with IITWares?+

Scope drives price, so we quote after a short discovery call rather than from a rate card. What we can share upfront is the range for comparable projects and exactly what is included, so the comparison against other proposals is fair.

Keep Reading

All Articles →

Ready to put these ideas to work?

Start a Project →