Housing Analyser.
Wiki · valuation

How value is determined.

The site reads value two independent ways — from comparable sales, and from a price model — then shows you both and whether they agree. Nothing here is a single magic number. Every figure carries its evidence, and the model's accuracy is measured against real sold outcomes, not asserted. Back to browse.

1 · Comps — what similar places actually sold for

€/m² of nearby sales

The most grounded read of value. For a listing, the site finds recent sales of comparable properties nearby — same property-type family (house / apartment / bungalow), within about one bedroom, within roughly ±15% of floor area, and in the same Eircode routing key (the first three characters, which is roughly a one-mile radius in Galway). If that yields too few, it widens the net rather than show nothing.

Each comp is inflation-adjusted to today's money using the CSO house-price index — a €200,000 Galway sale in 2014 is roughly €340,000 in 2026 money, and comparing without adjusting would make every older sale look cheap. The comps are then ranked by how close they are on distance, recency and size, so the most relevant ones surface first.

The headline comp figure is €/m² — price per square metre — taken as the median of the comp set (the middle value, not the average, so one odd penthouse can't swing it). The listing's own asking €/m² is compared against that median to produce the value score below.

Why thin data suppresses the score: only comps that actually carry a floor area contribute a €/m². If fewer than four do, the site shows no score at all rather than a confident- looking number resting on two sales. This is deliberate — it's the difference between “we don't know” and a false bargain.

2 · The AVM — a model that prices features

automated valuation model

The second, independent read. An AVM (automated valuation model) is a model trained on roughly 14,000 Galway sales to learn what each feature is worth: floor area, beds and baths, BER, property type, year- built bucket, and location. Hand it an active listing's features and it returns an estimated price.

It's a gradient-boosted tree model (LightGBM), chosen over a neural network because 14,000 rows is a small dataset where a tuned tree model gets within striking distance for a fraction of the complexity. It trains in seconds. Prices are learned in log space and inflation-adjusted, the same as the comps.

Crucially the AVM returns not just a point estimate but a band — a low (p10), middle (p50) and high (p90) estimate — so you see the model's uncertainty, not a single falsely-precise figure. The honest limits: it knows nothing about condition (a gutted comp and a renovated one look identical to it), and location signal is coarse, which is what the next step fixes.

3 · The residual correction — checking the model on your street

local bias adjustment

The AVM values a property from its features but has no sense of street or estate identity. Two houses with identical features on different roads get the same prediction, even when one estate consistently trades 8% above the model and the next trades below it. That gap is real, local, and visible in sales that already happened.

So before trusting the model's number, the site checks how wrong the model has recently been nearby: it looks at the model's error on recent sold properties close by, weights them by distance and recency, and nudges the prediction by that local bias. If the model has been running 6% low around here, the estimate shifts up about 6%. When few neighbours exist the correction shrinks toward zero — it only trusts a pocket once enough sales agree.

Measured on the backtest, this moved the model's median error from 10.14% to 9.89% — a modest, broad gain, biggest where the plain model was weakest (rural and county stock, where estate premium is strongest and the model's location signal is coarsest).

4 · Calibrated bands — an 80% band that's actually 80%

conformal calibration

A price band is only useful if it means what it says. An 80% band should contain the real sold price 8 times in 10 — and the point is that we measure whether it does. The raw AVM band was too tight: its nominal 80% band only actually covered about 70% of sold prices, so “looks under-priced” flags fired on noise.

The fix is conformal calibration: the band is widened by a data-driven margin, learned from how often recent sold prices fell outside it, until its measured coverage matches the 80% target. Because the model works in log space, the widening is proportional — it stretches a €150k cottage's band and a €1.2M house's band by the same relative amount, not a flat euro figure.

The honest cost: a truthful 80% band on this data is wide — roughly 48% of the price, versus the old tight (but dishonest) 39%. The old band bought its tightness by under-covering. One known gap remains: the cheap / rural under-€250k tail is still slightly under-covered, because it's genuinely the hardest segment to value.

5 · The production model — the two fixes, composed

what the site actually serves

The model the site serves is the AVM with both corrections composed together: the residual correction re-centres the point on local reality, and conformal calibration widens the band around that corrected point so its coverage is honest. They're composed carefully so the band isn't shifted twice.

These are the real shipped numbers, from a leak-free walk-forward backtest — the model is only ever scored on sales it was trained strictly before:

Median error (MdAPE)
9.8%
Within 10%
50.9%
Band coverage
79.5%

In plain terms: half its estimates land within 9.8% of the real sold price, just over half land within 10%, and its 80% band genuinely contains the sold price about 79.5% of the time — measured, not claimed. For comparison, the raw model before both fixes sat at 10.14% median error and only 69.6% band coverage.

The live card on the market trends page shows the current model's numbers and a scatter of its recent predictions against actual sold prices. The model retrains monthly and only swaps in if the new version passes a sanity check.

Model quality

The composed valuation model (residual-corrected point + conformal band, ADR 0057). Headline numbers are a leak-free walk-forward on real sold outcomes; the scatter scores the live model on recent sales. How this works →

no model on disk
MdAPE
Within 10%
Band coverage

No shipped validation metrics yet — train the composed model.

No recent sold rows the model can score yet.

Scatter: 0 most-recent sold rows the model could value.

6 · Value score vs the model — and why we show both

two methods, not one

The value score you see on a listing is the comps read: how far the asking €/m² sits below (or above) the median €/m² of comparable sales. A positive score means asking is below the comp median — potentially under-priced; negative means above it. It's deliberately simple and interpretable: “below half of similar properties' price per square metre” is plain English.

The AVM is the separate model read. The two are computed independently — comps from raw nearby sales, the model from learned feature prices — so when they agree, that's genuine corroboration. When they disagree sharply, the listing carries a “comps and the model disagree” warning telling you to dig in, rather than picking one and hiding the other. Showing both, and their agreement, is the whole design.

Confidence is tied to the real evidence: the count of comps that actually carry a €/m² (8 or more is high, 5–7 medium, fewer is low), and a wide spread in the comp set caps confidence down. The score is never reported without that sample size attached.

7 · The deal flag — and the lemon trap

good deal vs discounted-for-a-reason

A listing is flagged as looking under-priced when its asking price sits below the model's calibrated low estimate (p10) — i.e. cheaper than the model's 80% band, now that the band is honestly wide. But a low price has two very different explanations, and the site refuses to conflate them:

  • A genuine deal — a motivated seller, an off-season listing, a thin market for that exact property.
  • A discounted lemon — a poor BER, 110 days on the market, a third price drop. Correctly priced for its problems, not under-priced.

So a flag is a prompt, never a verdict. Every listing shows the residual signals next to the score — BER, days on market, the count and size of price drops — so a 12% discount “explained” by an F-rated BER and 110 days unsold reads as exactly what it is. The human (you) decides.

What we tried and rejected

honesty about dead ends

A repeat-sales index (the Case-Shiller approach) was considered and parked. The idea is appealing: track the same property selling twice over time and you get a pure price- change signal, free of any guess about which features matter. But it doesn't fit this data. Galway has only about 38,000 PPR sales across 15 years, and the subset of properties that sold more than once is far too sparse to estimate anything stable at the property level.

It also answers the wrong question — a repeat-sales method produces a market index, not a per-property valuation, and would need reliable same-property linkage across messy PPR address strings that we don't have. So the time-adjustment uses the official CSO index instead, and the per-property read stays with comps and the hedonic model above.

Two other extensions are deliberately deferred for the same KISS reason: a learned blend over comps + model (a second model to train and explain, for an unproven gain over the fixed correction that already works), and per- segment band margins (which would fix the under-€250k coverage gap but need far more calibration data). Both wait until a real decision is blocked without them.