Wards, MSOAs and why the geometries matter
data-democracy
data-analysis
place-based-change
Why we analyse local data by census geographies (OAs, LSOAs, MSOAs) rather than wards
Author

Celestin Okoroji

Published

September 29, 2026

Wards, MSOAs and why the geometries matter

Celestin Okoroji - 29th September 2026

One of the key things we often need to explain, when we are doing our Data Democracy work with communities, is the geographic unit of analysis. Communities are used to seeing things in terms of ‘wards’ but we almost exclusively use census units. The ones that end in OA (MSOA, LSOA, OA).

Once explained, people seem to both understand and appreciate why we don’t use wards in our analysis. This blog post shares that explanation with a wider audience and, in the Just Knowledge tradition, creates an ongoing resource for anyone to use.

TL;DR: Census units (the ones that end in OA) are easier to compare across place and time. Councils have good reasons to reach for wards, but those same reasons make wards an awkward unit for analysis.

To explain our approach, I’ll first set out how census geographies work. Then I’ll explain what wards are for, why the difference matters, and why councils use wards anyway.

Census geographies are built to be compared

The ONS builds its small-area geographies in three layers:

  • Output Areas (OAs): 100 to 625 people.
  • Lower Layer Super Output Areas (LSOAs): 1,000 to 3,000 people.
  • Middle Layer Super Output Areas (MSOAs): 5,000 to 15,000 people.

There are two features of these oddly named boundaries that are important to consider.

  1. Each OA sits inside exactly one LSOA, and each LSOA inside exactly one MSOA, and every MSOA fits inside exactly one local authority. The whole of England and Wales is covered this way.
  2. Each of the geographies has population limits by design: no MSOA has more than 15,000 people, and no OA has fewer than 100 (with very few exceptions).

To see how that works, scroll through one corner of Southwark.

This is one Output Area (OA), an OA, near Peckham Rye station. At the 2021 Census, people lived here.

It has neighbours, counting itself. Each is an Output Area (OA) of a similar size.

Together they make exactly one Lower Layer Super Output Area (LSOA): , home to people.

That LSOA sits alongside its neighbours. There are of them here.

Together they make exactly one Middle Layer Super Output Area (MSOA): , with residents.

Zoom out and there are MSOAs like it.

And together they make exactly one local authority: . people, LSOAs, OAs.

Source: Office for National Statistics, Output Areas (December 2021) boundaries and Census 2021. Contains OS data © Crown copyright and database right 2021.

The design allows for comparison. Because every census unit holds a broadly similar number of people, a rate in one MSOA (or LSOA/OA) means roughly the same thing as a rate in another. A 20% figure in Hackney and a 20% figure in Everton rest on similar-sized bases. Wards do not have this feature, as illustrated below.

The first chart plots every census unit and every ward in England and Wales by how many people live there. Each type of census unit forms a tight band: nine in ten MSOAs hold between about 5,800 and 11,600 people. Wards, on the other hand, smear across half the chart, from around 1,000 people to more than 30,000. Also notice that, by design, every census unit is bigger than the units inside it: an MSOA, no matter how small, must contain more than one LSOA, and an LSOA more than one OA.

Every Output Area, LSOA, MSOA and ward in England and Wales. Each dot is one area.

The second chart puts every type on the same footing. Each row has 100 dots, each standing for 1% of that type of area, placed by how big it is compared with a typical area of the same type. The census units pile up close to the typical size; wards run from about a fifth of a typical ward to three and a half times one.

Each dot is 1% of that type of area, placed by its size compared with the median area of the same type.

Source: Office for National Statistics, Census 2021 usual residents. Wards are December 2022 boundaries, the latest ONS publishes Census 2021 figures for.

The ONS also keeps these areas reasonably stable between censuses. Splits or merges only happen when populations have grown or shrunk past the limits above, so most areas survive from one census to the next. You can look at an LSOA in 2011 and the same LSOA in 2021 and know that any change happened in the place itself (though those familiar with census units will understand the joys of doing analysis across vintages where there are shifts).

Nevertheless, nothing is perfect. Boundaries rarely match how residents think about their own neighbourhood (few people say they live in “Southwark 023”) and we are actively developing work on how to create new units with the best of what the census has to offer and recognisable to people as their neighbourhoods (more on that soon!). But for asking how places differ, and how they change, they are the best tool we have.

Wards are built for representation

Wards do a different job. The Local Government Boundary Commission for England draws them so that each councillor represents roughly the same number of electors. It also weighs community identity in its decision-making.

Those are democratic aims, and they produce units of uneven size. A ward can elect one, two or three councillors, so a three-member ward can hold around three times the electorate of a single-member ward in the same council. Between councils the spread is wider still. And the Commission reviews boundaries every so often, which means the geometries change.

As such, wards are a democratic unit and aren’t really set up to be a statistical one.

Why the difference matters

Three problems follow from using wards for analysis.

Size. In a small ward, a handful of cases can swing a rate up or down, so you end up reading noise. In a large ward the opposite happens: the average smooths over everything underneath it. Any large unit does this, and wards are often the largest. A ward of 15,000 people is a small town. It can post a middling deprivation score while (as local people always know) some streets inside it sit among the most deprived in the country.

The people on those streets vanish from the picture, and they are often the people who most need the council to see them.

Time. When a boundary review redraws the wards, the series breaks. A ward’s employment rate in 2019 and its rate in 2023 may describe two different places that happen to share a name. Southwark went through exactly this in 2016, affecting the 2018 elections.

The geometries themselves. Geographers call this the modifiable areal unit problem: draw the boundaries differently and the same underlying data tells a different story. Every map of areas makes an argument about appropriate subdivisions, whether or not its author meant to.

Code For Nerds: building the data
library(dplyr)
library(readr)
library(sf)
library(jsonlite)

lad_code <- "E07000138" # Lincoln
cache <- file.path("data", "imd-units") # downloads are cached here
arcgis <- "https://services1.arcgis.com/ESMARspQHYMw9BZ9/arcgis/rest/services"
dir.create(cache, recursive = TRUE, showWarnings = FALSE)

# Download once, then read from the cache
cached_csv <- function(path, fetch) {
    if (!file.exists(path)) write_csv(fetch(), path)
    read_csv(path, show_col_types = FALSE)
}

# Page through an ONS lookup table, 1,000 rows at a time
arcgis_table <- function(service, fields) {
    pages <- list()
    offset <- 0
    repeat {
        j <- fromJSON(paste0(
            arcgis, "/", service, "/FeatureServer/0/query?where=1%3D1&outFields=", fields,
            "&returnGeometry=false&orderByFields=OA21CD&resultRecordCount=1000",
            "&resultOffset=", format(offset, scientific = FALSE), "&f=json"
        ))
        stopifnot(is.null(j$error))
        pages[[length(pages) + 1]] <- j$features$attributes
        offset <- offset + nrow(j$features$attributes)
        if (!isTRUE(j$exceededTransferLimit)) break
    }
    bind_rows(pages)
}

# Census 2021 usual residents for every OA, from the ONS Census API in batches
census_population <- function(codes) {
    h <- curl::new_handle(useragent = "curl/8.7.1")
    batches <- split(codes, ceiling(seq_along(codes) / 700))
    bind_rows(lapply(batches, function(b) {
        url <- paste0(
            "https://api.beta.ons.gov.uk/v1/population-types/UR/census-observations",
            "?dimensions=sex&area-type=oa,", paste(b, collapse = ",")
        )
        obs <- fromJSON(rawToChar(curl::curl_fetch_memory(url, handle = h)$content))$observations
        tibble(oa = vapply(obs$dimensions, \(d) d$option_id[1], ""), population = obs$observation) |>
            summarise(population = sum(population), .by = oa)
    }))
}

# ---- Inputs ---------------------------------------------------------------------
# IoD 2025 scores for every LSOA in England
iod <- cached_csv(file.path("data", "iod25.csv"), \() read_csv(paste0(
    "https://assets.publishing.service.gov.uk/media/68ff5daabcb10f6bf9bef911/",
    "File_7_IoD2025_All_Ranks_Scores_Deciles_Population_Denominators.csv"
))) |>
    select(lsoa = 1, lsoa_name = 2, score = 5)

# Which LSOA, MSOA and (May 2025) ward each Output Area sits in
oa_msoa <- cached_csv(file.path(cache, "oa_lsoa_msoa.csv"), \() arcgis_table(
    "OA_LSOA_MSOA_EW_DEC_2021_LU_v3", "OA21CD,LSOA21CD,MSOA21CD,MSOA21NM"
))
oa_ward <- cached_csv(file.path(cache, "oa_wd25.csv"), \() arcgis_table(
    "OA21_WD25_LAD25_EW_LU_v3", "OA21CD,WD25CD,WD25NM,LAD25CD,LAD25NM"
))
oa_pop <- cached_csv(file.path(cache, "oa_population.csv"), \() census_population(oa_msoa$OA21CD))

oa <- oa_msoa |>
    inner_join(oa_ward, by = "OA21CD") |>
    inner_join(oa_pop, by = c("OA21CD" = "oa")) |>
    inner_join(iod, by = c("LSOA21CD" = "lsoa")) |> # IoD covers England only
    select(
        oa = OA21CD, lsoa = LSOA21CD, lsoa_name, msoa = MSOA21CD, msoa_name = MSOA21NM,
        ward = WD25CD, ward_name = WD25NM, lad = LAD25CD, lad_name = LAD25NM,
        population, score
    )

# ---- Scores and deciles -----------------------------------------------------------
# Each OA carries its LSOA's score. A bigger unit's score is the population-weighted
# mean of its OAs. Deciles are cut so each holds a tenth of England's people, for
# LSOAs too, so all three units are ranked exactly the same way.
score_units <- function(unit) {
    oa |>
        summarise(
            score = weighted.mean(score, population), population = sum(population),
            .by = all_of(unit)
        ) |>
        arrange(desc(score)) |>
        mutate(decile = pmin(10, floor(10 * (cumsum(population) - population / 2) / sum(population)) + 1))
}
lsoa <- score_units("lsoa")
msoa <- score_units("msoa")
ward <- score_units("ward")

oa <- oa |>
    left_join(select(lsoa, lsoa, lsoa_decile = decile), by = "lsoa") |>
    left_join(select(msoa, msoa, msoa_decile = decile), by = "msoa") |>
    left_join(select(ward, ward, ward_decile = decile), by = "ward")

# The same split by number of wards instead: what share of people lands in "decile 1"?
ward_naive <- ward |> mutate(decile = ntile(-score, 10))
ward_d1_naive <- round(100 * sum(ward_naive$population[ward_naive$decile == 1]) / sum(ward$population))

# Share of each council's residents in England's most deprived tenth, by unit
council_share <- oa |>
    summarise(
        lsoa = 100 * sum(population[lsoa_decile == 1]) / sum(population),
        msoa = 100 * sum(population[msoa_decile == 1]) / sum(population),
        ward = 100 * sum(population[ward_decile == 1]) / sum(population),
        .by = c(lad, lad_name)
    ) |>
    mutate(across(lsoa:ward, round), across(lsoa:ward, \(x) min_rank(-x), .names = "rank_{.col}"))
lincoln_share <- as.list(filter(council_share, lad == lad_code))
council_rank <- \(name, unit) council_share[[paste0("rank_", unit)]][council_share$lad_name == name]
council_pct <- \(name, unit) council_share[[unit]][council_share$lad_name == name]
ordinal <- \(n) paste0(n, if (n %% 100 %in% 11:13) "th" else c("th", "st", "nd", "rd", rep("th", 6))[n %% 10 + 1])

# ---- Lincoln's boundaries -----------------------------------------------------------
# Full-resolution Output Areas, dissolved into LSOAs, MSOAs and wards so every
# layer shares exactly the same edges
lincoln <- filter(oa, lad == lad_code)
oa_bfc_path <- file.path(cache, paste0("oa_bfc_", lad_code, ".geojson"))
if (!file.exists(oa_bfc_path)) {
    lsoas <- unique(lincoln$lsoa)
    parts <- lapply(split(lsoas, ceiling(seq_along(lsoas) / 8)), function(codes) {
        where <- paste0("LSOA21CD IN ('", paste(codes, collapse = "','"), "')")
        st_read(paste0(
            arcgis, "/Output_Areas_2021_EW_BFC_V8/FeatureServer/0/query?where=",
            URLencode(where, reserved = TRUE), "&outFields=OA21CD&outSR=27700&f=geojson"
        ), quiet = TRUE)
    })
    st_write(bind_rows(parts), oa_bfc_path, quiet = TRUE)
}
lincoln_oa <- st_read(oa_bfc_path, quiet = TRUE) |>
    st_set_crs(27700) |>
    inner_join(lincoln, by = c("OA21CD" = "oa"))
stopifnot(nrow(lincoln_oa) == nrow(lincoln))

dissolve <- function(unit, name, decile) {
    lincoln_oa |>
        group_by(across(all_of(c(unit, name, decile)))) |>
        summarise(population = sum(population), .groups = "drop") |>
        select(code = all_of(unit), name = all_of(name), decile = all_of(decile), population) |>
        st_transform(4326)
}
layers <- list(
    lsoa = dissolve("lsoa", "lsoa_name", "lsoa_decile"),
    msoa = dissolve("msoa", "msoa_name", "msoa_decile"),
    ward = dissolve("ward", "ward_name", "ward_decile")
)

# ---- Write GeoJSON for the map (coordinates rounded to about 1 cm) -------------
feature_collection <- function(x) {
    x <- st_cast(x, "MULTIPOLYGON")
    list(type = "FeatureCollection", features = lapply(seq_len(nrow(x)), function(i) {
        list(
            type = "Feature",
            properties = list(
                code = x$code[i], name = x$name[i], decile = x$decile[i],
                population = x$population[i]
            ),
            geometry = list(type = "MultiPolygon", coordinates = lapply(
                st_geometry(x)[[i]], \(p) lapply(p, \(ring) unname(round(ring[, 1:2], 7)))
            ))
        )
    }))
}
b <- st_bbox(layers$lsoa)
write_json(
    list(
        place = "Lincoln",
        bounds = list(c(b[["ymin"]], b[["xmin"]]), c(b[["ymax"]], b[["xmax"]])),
        share = lincoln_share[c("lsoa", "msoa", "ward")],
        units = lapply(layers, feature_collection)
    ),
    file.path("..", "assets", "data", "imd-lincoln.json"),
    auto_unbox = TRUE, digits = NA
)

Lincoln shows how much this matters. The government publishes the Index of Multiple Deprivation 2025 for LSOAs. We averaged those scores up to MSOAs and to wards, in the same way1, and asked one question: what share of Lincoln’s residents live in England’s most deprived tenth?

By LSOA, 18%. By MSOA, 26%. By ward, 8%. Same place, same data, same method: one set of geometries puts a quarter of Lincoln in the most deprived tenth, the other fewer than one in ten. Switch between them on the map.

None of these percentages is ‘wrong’, and wards are not always the ones that hide things. Across England, wards and MSOAs each smooth away about as many people living in the most deprived LSOAs. The point is that the answer moves with the geometry. Lincoln is not alone: Hastings has 33% of its residents in the most deprived tenth by MSOA but 54% by ward, and Knowsley comes 1st of every council in England by MSOA but 11th by LSOA. Anything that hands out money or attention by “share of residents in the most deprived areas” (or something similar) inherits whichever geometries it was given.

Ward data has uses, but we need to be cautious and avoid slipping into comparing wards, since they are not built to be comparable. This is the most common issue we see in local data, maybe by some distance.

Why councils use wards anyway

So why do councils (and people familiar with council stuff) keep reaching for wards? Because councillors are elected to them.

A councillor answers to the people in their ward. Casework, ward forums, and the like, all run through it, and residents mostly know which ward they live in (few have heard of an MSOA, and fewer want to, at least at first). When a councillor asks what is happening in their area, the ward is the area they mean.

Recognition also matters. Our work in Northumberland Park is a case in point. There, the MSOA and the ward cover broadly the same ground and share a name. I suspect that overlap is part of why the work landed with so many different stakeholders: councillors and residents could recognise their place in the data straight away. Had the geometries and the names diverged, it might have been harder to get the community story into the right hands.

None of which means we should default to wards. The analysis should still run on units built for comparison, if a comparative story is what we want to tell. Nevertheless, there are ways to serve both needs using best-fit lookups. Alas, that is a story for another day.

A (tongue-in-cheek) field guide to the other geometries on maps

Wards aren’t the only geometries competing for your attention. Here is a short, partial and only slightly unfair guide to the rest.

Geography What it is How it relates to census units When to use it
Local authority Your council’s area. Every LSOA and MSOA sits inside exactly one. Often. But the Isles of Scilly (about 2,000 people) and Birmingham (over a million) are both “a local authority”, so compare with care.
Postcode (unit, sector, district, area) Royal Mail’s delivery system, in four levels: area (SE), district (SE15), sector (SE15 4) and the full unit postcode (SE15 4AB, around 15 addresses). A full postcode is granular enough to assign to an OA through ONS lookups. Sectors, districts and areas are their own thing. Full postcode as a lookup is a very reasonable choice. All other postcode levels as a unit of analysis: never (there’s a funny story about me saying exactly this in a big project; it didn’t end well).
Parliamentary constituency The area that elects an MP. Built from wards, so census units fit only by best-fit. When writing to your MP. Otherwise, treat it as a ward with ambitions.
Police Force Area The patch of one of 43 forces in England and Wales (British Transport Police don’t have a specific area). Built from local authorities, so census units nest inside. When the police data comes no smaller, which is often. The Met and Dyfed-Powys are both “a force”. Our R package policedatR works at this level (and at census-area levels). For stop and search, our dashboard takes a different cut: stop-and-search.justknowledge.org.uk
NHS areas (ICBs and friends) The NHS’s planning footprints. Mostly built from local authorities. Only if the NHS makes you. Thanks to perpetual redisorganisation2, they get redrawn more often than most people move house.
GP practice list The people registered with a practice. None. Patients live wherever they like. Never as a geography. A GP list is not a place.
Travel to Work Area An area where most people both live and work. Built from LSOAs. Labour market analysis. One of the few on this list that earns its keep.
Built-up area The ONS outline of towns and cities, traced from the buildings. Doesn’t nest. Follows the bricks. For “is this a town or a city?” arguments, which will outlive us all.
School catchment The area a school admits from. Doesn’t nest, and can shift year to year, sometimes street by street. Only if you enjoy correspondence with estate agents.
ITL regions (formerly NUTS) The UK’s official statistical regions, in three tiers. Built from local authorities. International comparisons: ITL mirrors the EU’s NUTS system, so UK regions line up with European and OECD regional statistics. Also fine for comparing UK regions with each other. Local analysis: no.

If in doubt, start with census units and translate outward.

The geometries decide who can be seen

At Just Knowledge we talk a lot about putting data into the hands of the communities it describes. Choosing the geography is part of that work. The unit you pick decides which places show up and which dissolve into an average. This choice is not a neutral matter.

So use census geographies to understand places, and wards to talk with the people accountable for them. Explore deprivation by MSOA in England on our map, and if you’re wrestling with ward data, get in touch.

Ultimately, choose your geometries on purpose. Somebody lives on the other side of every boundary.

Footnotes

  1. Each Output Area takes its LSOA’s score, and an MSOA or ward score is the average of its Output Areas, weighted by how many people live in each. Wards vary so much in size that if you simply split them into ten equal-sized groups, the “most deprived tenth” of wards holds 14% of England’s people. So for all three units we cut the deciles so that each holds a tenth of England’s people. For LSOAs, which are much closer in size, this moves only a few areas at the edges away from their official decile.↩︎

  2. We borrow “redisorganisation” from Oxman, Sackett, Chalmers and Prescott, A surrealistic mega-analysis of redisorganization theories, Journal of the Royal Society of Medicine 98 (2005): 563–568.↩︎

Stay Up to Date with Our Work