One brief × 6 messages × 5 languages. Find and fix workflow issues before the main assessment.
Understand the quality.
Improve every language.
Test whether Glossavia produces accurate, clear and usable screening posts from supplied documents—and how much human work is needed to make them ready.
Enter pilot program10 source briefs × 6 messages × 5 language versions
Start with one brief and a small batch.
The pilot is a separate practice area. Its posts are evaluated here and are never sent to Facebook.
- 1. Create your pilot campaignOpen Pilot source & drafts, paste the approved information and save it. A clinical reviewer or administrator confirms the source check.
- 2. Generate, then reviewChoose your languages and batch size. Review each complete post in the language queue, compare its English wording and save the evaluation.
- 3. See what is readyOpen Pilot & evaluation for language progress, reviewer names, scores and unresolved issues. Export the report for your team meeting.
Ten fresh briefs × 6 messages × 5 languages. Freeze the sample and preserve first-pass results.
20% of the main set, plus every major or critical finding. Compare ratings before discussing them.
Include different screening topics and document formats within the team’s intended scope.
English plus four languages chosen with local communities and competent reviewers.
Clinical source checking and language review of the complete caption and image.
100 is useful for finding early problems. Broader coverage needs more.
Start with a 30-post feasibility run (one brief × six messages × five languages) to repair workflow problems. Then evaluate a fresh, frozen set of 300 posts: 60 distinct message groups, with 60 assets per language. The feasibility run is separate from the main assessment.
This is a practical coverage target, not a statistical power calculation. Translations of the same message share a source and are related observations. Report results by source brief, message group and language. Even zero critical errors in this sample would not establish that the system is error-free.
From agreed protocol to a documented decision
- 01
Agree the sample and reviewers
Confirm the ten briefs, five language versions, intended audience and programme scope. Suggested language options include English, Tamil, Urdu, Hindi and Punjabi in the locally appropriate script. Each additional language needs its own review; it is not validated by results in another language.
Use sources covering eligibility where stated, informed choice, appointment changes, access or language support, and screening-versus-symptom distinctions where relevant. Include short or incomplete briefs and PDF-only links as separate challenge cases; the expected outcome may be a request for missing information.
- 02
Record a fair baseline
Prepare a matched subset manually using the same sources and languages. For a separate efficiency comparison, record staff drafting, translation, image preparation and review time on a matched subset. Timing is not required for each quality review in Glossavia. Record corrections in the single issue note. Capture source versions, generation date, model/settings and assessor competence before comparing workflows.
- 03
Generate complete posts
In each pilot campaign, save the source and choose “6 messages · balanced pilot sample”. Glossavia spreads these six message groups across 30 days, creates every selected language version and prepares the images before opening review. Other exploratory batches offer 1, 3, 10, 15, 20 or 30 messages per language. Keep exploratory checks separate from the agreed six-message main sample. No pilot batch sends posts to Facebook.
- 04
Assess every first version
Review source fidelity, language accuracy, clarity and informed choice, and image accuracy and legibility on a 1–5 scale. Check the complete caption, call to action, tags, short image headline and accessible image description. Record four scores and the highest issue severity in “Review & approval”. Add one note when an issue needs correction; the reviewer and date are recorded automatically. Clinical and language decisions are separate responsibilities.
- 05
Check agreement and correct errors
Independently double-rate a preselected 20% sample (60 assets: 12 per language spread across all ten briefs), plus every major or critical finding. Record ratings before discussing differences; adjudicate disagreements with an appropriate third reviewer. Keep the first-pass findings and assess corrected versions again.
- 06
Review with communities and decide
Use consented sessions with speakers of each language to explore comprehension, acceptability, cultural fit and clarity of the intended action. Record these findings alongside the technical review. Decide whether to revise the system, repeat affected tests or begin a limited service implementation.
Define errors before the first review
A credible risk of harm: wrong eligibility or clinical instruction, symptoms diverted into routine screening, or a materially false claim.
A substantial meaning, action, translation or visual error that prevents appropriate use and needs correction.
A small wording or layout issue that does not change meaning, choice or the intended action.
No identified error within the reviewer’s competence and the stated assessment criteria.
Score anchors: 1 unacceptable; 2 major revision; 3 some revision; 4 minor revision; 5 meets criterion. Severity and scores are recorded separately.
Proposed decision criteria
Use the following criteria for this pilot, and keep them fixed before reviewing the main sample. They are practical evaluation criteria, not proof of clinical effectiveness or an established validation standard.
- Every final asset has an exact supporting source, the required human checks and a complete legible image.
- No unresolved major or critical findings in assets considered for release. A critical first-pass error triggers investigation and reassessment of the affected generation method.
- A proposed first-pass target is at least 90% of assets in each language scoring 4 or 5 in all four domains, with no critical errors. Report actual counts and every correction, even if the target is met.
- Report generation failures, blocked briefs, first-pass error severity, review agreement, correction burden, staff time and API costs. Do not average away a weak language.
- Decide on time-saving targets after measuring the manual baseline. Include community comprehension and service workload in the decision.
What this pilot can establish
It can describe content quality, usability, review workload and failure patterns within the tested scope. It cannot prove clinical effectiveness, equitable reach or increased screening attendance. A later implementation study needs an agreed baseline, suitable comparison, measures of reach and engagement, and—where appropriate and authorised—attendance outcomes.
The wording approach draws on the Screening text message principles (5 April 2022). These are SMS principles adapted for clear, non-coercive social communication. Their patient-specific timings, 320-character preference and message-delivery rules are not social-media scheduling requirements.