Choose a program
PILOT PROGRAM · PROPOSED PROTOCOL

Understand the quality.
Improve every language.

Test whether Glossavia produces accurate, clear and usable screening posts from supplied documents—and how much human work is needed to make them ready.

Enter pilot program
300complete language-specific posts

10 source briefs × 6 messages × 5 language versions

YOUR FIRST SESSION

Start with one brief and a small batch.

The pilot is a separate practice area. Its posts are evaluated here and are never sent to Facebook.

  1. 1. Create your pilot campaignOpen Pilot source & drafts, paste the approved information and save it. A clinical reviewer or administrator confirms the source check.
  2. 2. Generate, then reviewChoose your languages and batch size. Review each complete post in the language queue, compare its English wording and save the evaluation.
  3. 3. See what is readyOpen Pilot & evaluation for language progress, reviewer names, scores and unresolved issues. Export the report for your team meeting.
Open pilot workspace
30First, a feasibility run

One brief × 6 messages × 5 languages. Find and fix workflow issues before the main assessment.

300Then, the main assessment

Ten fresh briefs × 6 messages × 5 languages. Freeze the sample and preserve first-pass results.

60Independent second reviews

20% of the main set, plus every major or critical finding. Compare ratings before discussing them.

10 different briefs

Include different screening topics and document formats within the team’s intended scope.

5 language versions

English plus four languages chosen with local communities and competent reviewers.

Two perspectives

Clinical source checking and language review of the complete caption and image.

WHY 300?

100 is useful for finding early problems. Broader coverage needs more.

Start with a 30-post feasibility run (one brief × six messages × five languages) to repair workflow problems. Then evaluate a fresh, frozen set of 300 posts: 60 distinct message groups, with 60 assets per language. The feasibility run is separate from the main assessment.

This is a practical coverage target, not a statistical power calculation. Translations of the same message share a source and are related observations. Report results by source brief, message group and language. Even zero critical errors in this sample would not establish that the system is error-free.

THE PROCESS

From agreed protocol to a documented decision

  1. 01

    Agree the sample and reviewers

    Confirm the ten briefs, five language versions, intended audience and programme scope. Suggested language options include English, Tamil, Urdu, Hindi and Punjabi in the locally appropriate script. Each additional language needs its own review; it is not validated by results in another language.

    Use sources covering eligibility where stated, informed choice, appointment changes, access or language support, and screening-versus-symptom distinctions where relevant. Include short or incomplete briefs and PDF-only links as separate challenge cases; the expected outcome may be a request for missing information.

  2. 02

    Record a fair baseline

    Prepare a matched subset manually using the same sources and languages. For a separate efficiency comparison, record staff drafting, translation, image preparation and review time on a matched subset. Timing is not required for each quality review in Glossavia. Record corrections in the single issue note. Capture source versions, generation date, model/settings and assessor competence before comparing workflows.

  3. 03

    Generate complete posts

    In each pilot campaign, save the source and choose “6 messages · balanced pilot sample”. Glossavia spreads these six message groups across 30 days, creates every selected language version and prepares the images before opening review. Other exploratory batches offer 1, 3, 10, 15, 20 or 30 messages per language. Keep exploratory checks separate from the agreed six-message main sample. No pilot batch sends posts to Facebook.

  4. 04

    Assess every first version

    Review source fidelity, language accuracy, clarity and informed choice, and image accuracy and legibility on a 1–5 scale. Check the complete caption, call to action, tags, short image headline and accessible image description. Record four scores and the highest issue severity in “Review & approval”. Add one note when an issue needs correction; the reviewer and date are recorded automatically. Clinical and language decisions are separate responsibilities.

  5. 05

    Check agreement and correct errors

    Independently double-rate a preselected 20% sample (60 assets: 12 per language spread across all ten briefs), plus every major or critical finding. Record ratings before discussing differences; adjudicate disagreements with an appropriate third reviewer. Keep the first-pass findings and assess corrected versions again.

  6. 06

    Review with communities and decide

    Use consented sessions with speakers of each language to explore comprehension, acceptability, cultural fit and clarity of the intended action. Record these findings alongside the technical review. Decide whether to revise the system, repeat affected tests or begin a limited service implementation.

CONSISTENT ASSESSMENT

Define errors before the first review

Critical

A credible risk of harm: wrong eligibility or clinical instruction, symptoms diverted into routine screening, or a materially false claim.

Major

A substantial meaning, action, translation or visual error that prevents appropriate use and needs correction.

Minor

A small wording or layout issue that does not change meaning, choice or the intended action.

None

No identified error within the reviewer’s competence and the stated assessment criteria.

Score anchors: 1 unacceptable; 2 major revision; 3 some revision; 4 minor revision; 5 meets criterion. Severity and scores are recorded separately.

Proposed decision criteria

Use the following criteria for this pilot, and keep them fixed before reviewing the main sample. They are practical evaluation criteria, not proof of clinical effectiveness or an established validation standard.

  • Every final asset has an exact supporting source, the required human checks and a complete legible image.
  • No unresolved major or critical findings in assets considered for release. A critical first-pass error triggers investigation and reassessment of the affected generation method.
  • A proposed first-pass target is at least 90% of assets in each language scoring 4 or 5 in all four domains, with no critical errors. Report actual counts and every correction, even if the target is met.
  • Report generation failures, blocked briefs, first-pass error severity, review agreement, correction burden, staff time and API costs. Do not average away a weak language.
  • Decide on time-saving targets after measuring the manual baseline. Include community comprehension and service workload in the decision.

What this pilot can establish

It can describe content quality, usability, review workload and failure patterns within the tested scope. It cannot prove clinical effectiveness, equitable reach or increased screening attendance. A later implementation study needs an agreed baseline, suitable comparison, measures of reach and engagement, and—where appropriate and authorised—attendance outcomes.

The wording approach draws on the Screening text message principles (5 April 2022). These are SMS principles adapted for clear, non-coercive social communication. Their patient-specific timings, 320-character preference and message-delivery rules are not social-media scheduling requirements.