Research Protocols

100 Saudi Apps: An RTL UX Benchmark — Research Protocol

Proposed methodology for 100 Saudi Apps. Study design, sampling, evidence collection and limitations. No completed benchmark results.

··11 min read
On this page
  1. Executive summary
  2. Research question and scope
  3. Define the sampling frame before collecting examples
  4. Build a coding framework
  5. Evidence capture and audit trail
  6. Analysis plan
  7. Results section template — fill only after data collection
  8. How to turn the benchmark into an SEO, GEO, and citation asset
  9. Visuals to produce
  10. Quality checks before publication
  11. CTA
  12. FAQ

Executive summary

100 Saudi Apps: An RTL UX Benchmark can become a high-value original research asset for os3li.com because it creates evidence that other designers, founders, product teams, journalists, and AI answer systems can cite. The value, however, comes from reproducibility rather than from a dramatic headline. This protocol defines how to choose the sample, what to observe, how to code each product consistently, how to separate descriptive findings from interpretation, and how to publish the final benchmark without inventing precision. The article should be updated on a visible cadence so readers can distinguish a current snapshot from a permanent claim. Until the study is run, the sections marked Results should remain unpublished or clearly labeled as pending.

Research question and scope

The primary research question is: what observable design patterns and failure modes appear across the sample relevant to 100 Saudi Apps?. A secondary question should examine how those patterns vary by market, language mode, platform, or product category when the sample is large enough to support the comparison. Do not start with a conclusion such as 'most apps fail at RTL.' Start with criteria that can be observed and coded consistently. The study should be descriptive unless you also collect behavioral evidence; screenshots can show that a pattern exists, but they cannot prove that the pattern caused conversion, satisfaction, or retention.

Define the sampling frame before collecting examples

Sampling rule 1

Write an inclusion rule that another researcher could follow without asking you which apps you 'meant'. The purpose is comparability: readers should be able to understand what the benchmark represents and what it does not represent. If access changes during collection, log the change rather than replacing the item silently.

Sampling rule 2

Record the country or market, category, platform, app or web version, language state, and capture date for every item. The purpose is comparability: readers should be able to understand what the benchmark represents and what it does not represent. If access changes during collection, log the change rather than replacing the item silently.

Sampling rule 3

Decide whether the unit of analysis is a whole product, a single flow, a screen, or a component. Keep it consistent. The purpose is comparability: readers should be able to understand what the benchmark represents and what it does not represent. If access changes during collection, log the change rather than replacing the item silently.

Sampling rule 4

Avoid selecting only famous or visually polished products; that creates a prestige sample rather than a market sample. The purpose is comparability: readers should be able to understand what the benchmark represents and what it does not represent. If access changes during collection, log the change rather than replacing the item silently.

Sampling rule 5

If the title contains a number such as 100 apps, publish the exact list or a defensible anonymized list and explain any exclusions. The purpose is comparability: readers should be able to understand what the benchmark represents and what it does not represent. If access changes during collection, log the change rather than replacing the item silently.

Sampling rule 6

Freeze the sample for the main analysis. New examples can be added in a later edition rather than quietly changing the denominator. The purpose is comparability: readers should be able to understand what the benchmark represents and what it does not represent. If access changes during collection, log the change rather than replacing the item silently.

Build a coding framework

Criterion 1: Mirroring rules

For mirroring rules, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Criterion 2: Reading order

For reading order, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Criterion 3: Bidirectional content

For bidirectional content, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Criterion 4: Rtl navigation

For RTL navigation, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Criterion 5: Direction-aware components

For direction-aware components, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Criterion 6: Arabic typography

For Arabic typography, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Criterion 7: Localization qa

For localization QA, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Criterion 8: Research question

For research question, define an observable yes/no, categorical, or ordinal code before reviewing the full sample. Write examples of what counts and what does not count in the codebook. For 100 Saudi Apps: An RTL UX Benchmark, ambiguity here is more dangerous than having a smaller number of criteria. Where the judgement is subjective, have a second reviewer code a subset independently and compare disagreements. Keep a notes field for edge cases, but do not change the code definition simply because one product is unusual. If the criterion cannot be observed from the available interface, record it as not observable rather than guessing.

Evidence capture and audit trail

For every coded observation, store enough evidence to revisit the decision: screenshot or recording, product/version identifier, locale, date, step in the flow, and a short note. Use stable file naming so the article can link a chart back to its source evidence. If a product contains personal data, do not capture or publish it; create a clean test account or redact sensitive information. For changing interfaces, timestamps matter because a later version may fix or introduce the pattern. A research benchmark becomes more credible when readers can inspect the methodology even if the raw dataset cannot be fully public.

Analysis plan

  1. Start with descriptive counts and proportions for each criterion.

  2. Compare meaningful segments only when each segment has enough observations to avoid misleading percentages.

  3. List recurring combinations of patterns, not just isolated frequencies.

  4. Separate product-category differences from country or language differences where possible.

  5. Investigate outliers because they often reveal an alternative pattern the codebook did not anticipate.

  6. Do not infer user preference or business impact from prevalence alone.

  7. Write the limitations before writing the conclusion; this reduces the temptation to overstate the story.

Results section template — fill only after data collection

Use this structure after the data exists:

  1. Sample: what was reviewed, when, and under what inclusion rules.
  2. Top-level findings: three to five patterns supported by the data.
  3. Breakdown by criterion: chart/table plus examples and counterexamples.
  4. Breakdown by market/category: only when the sample supports it.
  5. Notable outliers: useful alternatives that challenge the dominant pattern.
  6. Design implications: what a product team should test, not what the data magically “proves.”
  7. Limitations: sampling, access, observational limits, changing product versions.
  8. Dataset/method link: enough material for a reader to reproduce the analysis.

Never fill a missing number with an estimate because it “sounds realistic.” If data is incomplete, say so.

How to turn the benchmark into an SEO, GEO, and citation asset

Original research can perform well in classic search and AI-generated answers because it offers a sourceable fact pattern rather than another summary of common advice. Make each chart understandable without the surrounding paragraph, give tables descriptive headings, and write a one-sentence finding directly above or below the evidence. Create a stable canonical URL, keep the previous edition accessible if methodology changes materially, and show the date of data collection as well as the article update date. Use an author profile, methodology page, and internal links to the relevant Arabic/RTL, AI UX, research, or MENA pillar pages. When the study is refreshed, update the page rather than publishing near-duplicate versions unless the historical comparison itself is valuable.

Visuals to produce

  • A methodology diagram showing sample → capture → coding → review → analysis.

  • A sample composition chart by market/category/language.

  • One chart per major criterion with the denominator visible.

  • Annotated screenshots that illustrate a pattern without exposing private data.

  • A comparison matrix showing good alternatives rather than a simplistic best/worst ranking.

  • A downloadable CSV or methodology appendix when licensing and privacy allow it.

Quality checks before publication

  • Two people can apply the codebook to the same example and reach broadly consistent results.

  • Every headline statistic has a visible denominator and a traceable data source.

  • No causal language is used for purely observational evidence.

  • The article distinguishes MENA-wide observations from country-specific findings.

  • The sample and capture period are visible near the top of the article.

  • Screenshots are lawful to use and contain no personal or confidential information.

  • The CTA is separated from the research findings so commercial intent does not distort interpretation.

CTA

If your team wants to run a product benchmark, Arabic/RTL audit, UX research study, or competitive UX analysis with a reproducible methodology, contact Osama Ali or send a WhatsApp message.

FAQ

Can this article be published before the benchmark is complete?

Publish the methodology as a research plan if that is useful, but do not present it as completed research. The final benchmark title should only carry numerical or market-wide claims after the sample has been collected and coded. A transparent “study in progress” page is better than invented findings.

Do screenshots prove that a pattern is good or bad?

No. Screenshots can document implementation patterns. They cannot prove user preference, comprehension, conversion impact, or causal effects. Those claims require behavioral research, analytics, experiments, or other evidence.

How large should the sample be?

The number depends on the claim. A small sample can support an exploratory pattern library; a broad market claim needs a defensible sampling frame and enough coverage to avoid making a handful of products represent an entire region. Publish the limits explicitly.

How should MENA be handled?

Treat the region as a grouping, not a single user type. Record country, language state, and category. Compare countries only when the sample supports it, and avoid turning a product implementation choice into a cultural stereotype.

How does this help AI-search visibility?

Clear, sourceable findings, tables, definitions, and stable URLs give retrieval systems material they can ground and cite. That does not guarantee inclusion in any AI answer, but it makes the content easier to interpret and verify.

What should be updated in future editions?

Refresh the sample or rerun the same sample, preserve the methodology, note product/version changes, and publish a comparison against the previous edition. Do not silently rewrite old numbers without a change log.