---
title: "What 575 Sri Lankan government websites tell us, and what AI got right"
author: Amila Sampath
date: 2026-10-05
updated: 2026-10-07
url: https://www.govlk.site/blog/what-575-government-websites-tell-us/
tags: ["GovTech", "Sri Lanka", "UX research", "Accessibility", "Data visualisation", "AI"]
description: "Amila Sampath built Overlook on top of Lanka Data Foundation's volunteer scorecard to audit 575 Sri Lankan government websites with a browser robot and AI. Phone speed is the biggest problem, and after three scans the automated scorecard matched the volunteers' grade on 75.4% of sites."
---

# What 575 Sri Lankan government websites tell us, and what AI got right

By Amila Sampath, published 5 October 2026. Independent experiment, not an official audit.

Lanka Data Foundation's volunteers scored every Sri Lankan government website in one day. I tried to rebuild that research alone, with a browser robot and an AI assistant. Here is what I built, what it found, and where the machine and the volunteers disagreed.

## In short

- Overlook is an independent experiment that captures and audits all 575 Sri Lankan government websites in the Lanka Data Foundation (LDF) scorecard, on desktop and on a phone.
- Phone speed is the biggest problem: the median Lighthouse mobile performance score is 32 out of 100 across 494 tested sites, and only 7 sites show their main content within 2.5 seconds.
- Only 7 sites clear every basic: online, valid SSL, gov.lk domain, a language switcher, accessibility 90+ and phone speed 50+.
- An automated scan of LDF's rubric matched the volunteers' grade on 75.4% of the 513 sites both could load, and within one grade on 93.2% (correlation 0.85).
- Of 381 disagreements, 222 were judgement calls in the rubric, 72 were volunteer errors and 27 were scan errors.

![The Overlook board: a grid of Sri Lankan government home pages on the left, sorted by LDF score, with government design system reference pages from the UK, US, Canada, Australia, NSW and Estonia on the right. Header stats read 556 of 575 sites shown, average LDF score 69, median Lighthouse performance 32 and accessibility 85.](https://www.govlk.site/blog/what-575-government-websites-tell-us/img/overlook-wall.jpg)

*Overlook, the tool this post is about. Every government home page on one zoomable wall, with other governments' design systems in a column on the right. [Open it at www.govlk.site](https://www.govlk.site).*

## Where this started

In September, [Lanka Data Foundation](https://opendata.lk/) (LDF) ran its Government Website Scorecard Digital Audit. On 5 September 2026, 24 volunteers sat down together and checked 577 government websites against a simple, public rubric: does the site load, does it have a valid SSL certificate, can you read it in Sinhala, Tamil and English, does it use the official gov.lk domain, and does its contact email match that domain. They published the results on 25 September.

I loved it. It is the kind of civic research that rarely happens because it is slow and boring to do by hand. It also left me with a question I could not shake: how much of this could one person do alone with today's AI tools, and could they go further than the rubric?

So I tried. The first version took me about eight hours. Everything since has been about checking whether the numbers are right.

- **575**: sites in the LDF scorecard list I worked from
- **24**: LDF volunteers, one day, by hand
- **1**: person, about 8 hours for the first build
- **34,769**: screenshots and crops on the live board

## What I built: Overlook

Overlook takes LDF's list of sites and opens every one in a real browser, first as a desktop at 1280 × 800, then as a phone at 390 × 844. On each site it:

- takes the first screen as a citizen lands on it, clicks through any "choose your language" page to the English home page, then visits one inner page;
- finds the common parts of the page (header, menu, search, hero, footer, logo, page layout) and crops each one out of the live page;
- runs Lighthouse for speed on a phone, plus axe-core and HTML_CodeSniffer for accessibility, and its own checks for Sinhala and Tamil language tags, tap targets and skip links;
- tries the components the way a person would: types "contact" into search and reads what comes back, opens the phone menu, tabs through the main menu with a keyboard, and watches the hero slider for motion and a pause button.

Then it lays everything out side by side on one big zoomable wall. You can sort by LDF score, group by ministry or site type, and click any tile for that site's full sheet. A "See as" mode shows any page the way someone with colour blindness, low vision or cataracts would see it.

Next to our sites sits a column of government design systems: GOV.UK, the US Web Design System, Canada, Australia's AgDS, NSW and Estonia's TEDI, rebuilt from their published code. When you look at 258 search boxes from our estate, you can see right away what a well-specified one looks like.

### Ask the data

The wall answers "what does it look like". People kept asking "so what does it mean". So I added a chat that sits on top of everything the tool measured. You ask in plain words, like "Which site types are slowest on a phone?", and get a short answer with a live chart. Click a bar and the wall opens on those exact sites.

Because this is public data about public institutions, I spent most of my time on guard rails rather than features. Every number in an answer is checked against what the database actually returned before you see it, and any figure the AI cannot back up is removed. Charts are re-run with the data tables emptied, so a number typed in by the model, rather than counted, gets caught. It costs well under one US cent per question.

[Video: A 60-second walkthrough of asking Overlook a question and opening a shared dashboard](https://www.govlk.site/blog/what-575-government-websites-tell-us/img/walkthrough.mp4)

*A 60-second walkthrough: ask a question, get a chart, open the sites behind it, turn it into a shareable dashboard.*

## What it found

The LDF rubric is about basics, and the estate does well on some of them. 95.1% of online sites have a valid SSL certificate. Languages are where the points are lost: only 51.2% offer a language switcher. Once I looked past the rubric, at what a citizen actually experiences on a phone, the picture changed a lot.

### Speed on a phone is the biggest problem

Of the 494 sites Lighthouse could test, the median mobile performance score is 32 out of 100. Only 5 score 90 or more. On a simulated mid-range phone, a typical home page takes about 15 seconds to show its main content, and only 7 sites do it within the 2.5-second "good" mark.

**Chart: Most government sites score under 50 for speed on a phone.** Lighthouse 13 mobile performance score, 494 sites, grouped in tens. Captured 30 September 2026. Lighthouse mobile performance scores of 494 sites, grouped in tens. Median 32; 5 sites score 90 or more.

A good scorecard grade does not mean a fast site: 84% of grade-A sites still score under 50 on phone performance. The scorecard and the experience measure different things, and both matter.

### The same component, built many different ways

- 258 home-page search boxes come in 181 different sizes. Search shows up in 7 different forms, and 224 of 519 home pages have no search at all.
- 53 of 490 phone menu buttons did not open when tapped.
- 124 home-page carousels move on their own. Only 1 has a pause button.
- Only 16.6% of sites have a skip link, something every design system gives you ready-made.

### Languages and accessibility

- 12% of working sites open on a "choose your language" gate before you see any content.
- 93% of pages with Sinhala or Tamil text do not tag that text as Sinhala or Tamil, so screen readers read it with the wrong voice, or not at all.
- The median Lighthouse accessibility score is 85.

### Security and housekeeping

- 260 of 545 reachable sites send none of the six standard security headers.
- 230 sites load trackers before the visitor has agreed to cookies.
- Home pages weigh 7.7 MB on average.

> Only 7 sites clear every basic: online, valid SSL, gov.lk, a language switcher, accessibility 90+ and phone speed 50+.

## The real test: would the machine agree with the volunteers?

Pictures and Lighthouse scores are one thing. The harder question was whether automation could reproduce LDF's human scorecard on the same rubric. If it could, the scorecard could be re-run every month instead of once a year. If it could not, I wanted to know exactly where and why.

I wrote a crawler that opens each site's core pages (Home, About, Services, Contact, RTI, News), switches into Sinhala and Tamil, counts which script each page is really written in, checks the SSL certificate on every address a citizen might type, reads the contact emails and applies the domain rules from LDF's report. Then I compared its score with the volunteers' score for every site.

### The first attempt was not good

On the first run, the scan matched the volunteers' grade on only 35% of sites. My first instinct was to blame the volunteers. When I looked closer, most of the gap was my own scan's fault. It counted any Sinhala word in the page code as a language switcher, credited full translation when a switcher merely offered a language without opening it, judged every email on six pages when volunteers judge the main one, and checked the certificate on the bare domain when many certificates only cover the www address.

So I rebuilt it to behave more like a volunteer: only count language controls a person can see, actually open each language version, and treat a machine translation widget as "dynamic" the way the rubric does. Then I did it again, adding language menus that open on click, image buttons, splash pages and bot walls.

**Chart: Three scans, closer each time.** Share of sites where the scan's grade matches the volunteers' grade. Grade agreement between the volunteers and the automated scan across three scan versions. Exact grade match rose from 35.3% to 75.4%, within one grade from 70.4% to 93.2%.

**Chart: Agreement by rubric item, first scan to final scan.** Share of sites where the scan and the volunteers gave the same points. Share of sites where volunteers and the scan agree, per rubric item, first scan versus final scan.

*Scans 1 and 2 cover all 575 listed sites; scan 3 covers the 513 sites both the volunteers and the scan could load, which is the fair comparison. r is the correlation between the two scores.*

On the final scan, across the 513 sites both sides could load, the grades match exactly on 75.4% of sites and are within one grade on 93.2%. The average score is 72.8 for the scan and 72.9 for the volunteers, the correlation is 0.85, and the typical gap is 5.4 points. Only 6 sites are three or more grades apart.

**Chart: Volunteer score against automated score, per site.** 513 sites both could load. Dots on the dashed line got the same score from both. Hover a dot for the site. Volunteer score against automated score for 513 sites both could load. Correlation 0.85.

### Who was right when they disagreed?

Agreement numbers hide the interesting part, so I went through every disagreement. For the ones the evidence could not settle, I took screenshots with images on, clicked the language controls by hand and traced the network requests. There were 381 category-level disagreements across 274 sites.

**Chart: Most disagreements are judgement calls, not mistakes.** Verdicts on 381 category-level disagreements, after manual review on 5 October 2026. Verdicts on 381 category-level disagreements between volunteers and the scan.

More than half are judgement calls: the rubric does not say exactly which pages count as "core pages" or how much Sinhala makes a page "translated", and volunteers scored identical platforms differently. The volunteers were wrong more often than the scan (72 against 27). A few examples:

- 12 sites from one ministry's section were published 20 points below the sum of their own criteria, because the translation points were dropped. irrigation.gov.lk shows 79 where its own criteria add up to 99.
- 24 sites are marked "email matches the domain" while their contact address is a Gmail or other-domain address.
- 21 disagreements are simply sites that changed between 5 September and our scan.

This is not a criticism of the volunteers. Twenty-four people scoring 577 sites in one day will make slips, which is exactly why a second, automated pass is useful. The point is that the two work best together.

> Automation is reliable on the deterministic checks. It struggles where a person has to read: is this page really in Tamil?

SSL, domain and switcher checks now agree on 95 to 97.5% of sites. The weakest item is still the count of core pages in all three languages, at 59.3%, and that is where the rubric itself needs a tighter definition before any method, human or machine, can be consistent.

## Three dashboards, three audiences

Once the data was trustworthy, the next question from people was "what should I do with it?". The chat can turn any answer into a shareable dashboard with one click. No login, and everyone sees the same charts. I made three ready-made ones, each with 13 or 14 charts, and every chart opens the real sites behind it.

- [**For designers**](https://www.govlk.site/#dash=897sks): What citizens meet on the home page: carousels, search, skip links, layouts and components compared with design systems.
- [**For developers**](https://www.govlk.site/#dash=d7zz62): Speed, page weight, security headers, platforms and CMS versions across the estate.
- [**For managers**](https://www.govlk.site/#dash=rsajkz): The whole estate at a glance: grades by site type, compliance, and a risk matrix of what to fix first.

## Why this matters for GovTech

Sri Lanka's digital government work is moving to [GovTech Sri Lanka](https://govtech.lk/), the state company the Cabinet approved in 2025 to [take over ICTA's role](https://www.ft.lk/front-page/GovTech-to-take-over-ICTA-role/44-776948) and lead national digital transformation. Projects like GovPay and a single citizen entry point depend on government websites people can actually use: on a phone, in their own language, and safely.

I think this experiment is useful to that work in four ways:

- **A baseline.** The estate now has a measured starting point beyond the rubric: speed, accessibility, security headers, components and languages, dated and repeatable.
- **Fixes that help many sites at once.** Many problems come from shared templates and the same few platforms. Fix the template and dozens of sites improve together.
- **A case for a shared design system.** 181 sizes of search box is what happens without one. Every system on the right-hand column of the wall shows a tested, accessible alternative, and the wall shows which patterns our sites already share.
- **A cheaper scorecard.** The rules-based checks agree with volunteers on 95% or more of sites. They can run every month, leaving people free for the parts that need human judgement.

### What this is, and what it is not

Overlook is an independent experiment. It is not an official audit and not a Government of Sri Lanka or LDF product. The LDF scores, grades and ranks are LDF's published numbers from 5 September 2026. Screenshots and components were captured on 29 September 2026, Lighthouse ran on 30 September (mobile, simulated throttling, performance moves about 5 to 10 points between runs), and the human-vs-AI comparison is from 5 October 2026. Sites change all the time, so any single number may be out of date for a given site. Chats on the board are recorded so I can improve the answers.

## What I learned

- **One person can now do quantitative research across hundreds of sites.** The building is fast. The checking is the real work, and it should take most of your time.
- **When the machine disagrees with people, look at your own method first.** My first scan was wrong far more often than the volunteers were.
- **Humans and automation are better together.** The scan caught volunteer slips; the volunteers set the standard the scan had to learn.
- **Showing beats telling.** A wall of 575 home pages makes the case for a shared design system faster than any table.

Thank you to Lanka Data Foundation and its volunteers for doing the hard work first. None of this would exist without their list and their scorecard. If you work on government websites, at GovTech, in a ministry or as a vendor, I would love to hear what would make this more useful to you.

## Links

- [**Overlook**](https://www.govlk.site): The board, the chat and the dashboards
- [**Lanka Data Foundation**](https://opendata.lk/): The volunteer scorecard this builds on
- [**GovTech Sri Lanka**](https://govtech.lk/): The government's digital transformation agency

## Questions people ask

### What is Overlook?

Overlook (www.govlk.site) is an independent experiment by Amila Sampath that puts every Sri Lankan government website from the Lanka Data Foundation scorecard on one zoomable wall, with component, speed, accessibility and language audits and an AI chat over the data. It is not an official government or LDF product.

### How fast are Sri Lankan government websites on a phone?

Across 494 sites tested with Lighthouse 13 on 30 September 2026, the median mobile performance score was 32 out of 100. Only 5 sites scored 90 or more, and only 7 showed their main content within 2.5 seconds.

### Can AI reproduce the LDF volunteer scorecard?

Largely. On the 513 sites both could load, the automated scan matched the volunteers' grade on 75.4% of sites and was within one grade on 93.2%. Agreement is 95% to 97.5% on SSL, domain and language switcher checks, and lowest (59.3%) on counting core pages in Sinhala, Tamil and English.

### Why does this matter for GovTech Sri Lanka?

It gives a dated, repeatable baseline of the government web estate, shows fixes that help many sites at once through shared templates, makes the case for a shared design system, and shows that the rules-based scorecard checks could run every month.

---

Data from the LDF scorecard (5 September 2026) and Overlook scans (29 September to 5 October 2026). Source page: https://www.govlk.site/blog/what-575-government-websites-tell-us/
