How These Numbers Are Made

Where every figure on this site comes from. Which ones a script re-checks on every change, which ones nothing watches, and how much weight to put on each.

Why this page exists

This site says Elmira's property taxes fall unfairly. That claim is only worth something if the numbers behind it are right.

Until now you had no way to test that. The explanations lived in code comments that nobody outside the project reads. So this page sets out how a figure is built, from the state record it starts as to the sentence you read. It also says plainly which figures are checked automatically and which are not, because that difference matters more than anything else here. Where the method itself is a judgment call, that is the other page: known limits and judgment calls.

This is not a list of sources. Sources live on the Data & Sources page, with links to the originals. This page is about what happens to that data after we download it.

Every source, and everywhere it ends up
The whole pipeline at once: what we read, what processes it, and which pages the result reaches. Hover or tab to any box to follow one path through.

Loading the source map…

A green outline means a script re-checks that page's figures on every change. An amber one means nothing does. Click any page to jump to its recipe below. This map is drawn from the same file the recipes come from, so it cannot drift from them.


Why the fork matters
A figure in a chart and the same figure in a sentence are not equally trustworthy, and the difference is not obvious from reading the page.
Drawn by machine

Charts look after themselves

When you open the assessment page, your browser downloads jcurve.json and draws the chart from whatever it finds there.

Re-run the script with newer sales, and every chart reading that file changes on its own. No one has to notice, remember, or retype anything.

Typed by a person

Sentences stop tracking the data

When someone writes "38.9% of the city's assessed value is off the tax rolls" into a paragraph, that figure is fixed. It was right the day it was typed.

If next year's roll says 41%, the chart beside it will move and the sentence will not. The page then contradicts its own graph, and nothing about it looks wrong.

Nearly every page here mixes both. So the question worth asking about any figure on this site is not "where did it come from" — that is in the footer of every page. It is "is anything still watching it?" For most of the site, the answer is no. The rest of this page is about that split.


Pages with a checking script
Nine pages are re-checked in full whenever anything changes: the home page, the four IDA pages plus the IDA board page, the zoning page, the assessment regressivity page, and the reassessment page.

For these, a change is checked by a script that actively tries to prove the page wrong. Here is what it does.

The fast checks now run themselves. Every proposed change is checked automatically, before anyone reviews it. That covers the home page, the five IDA pages, and a sweep for figures we have withdrawn and replaced. It needs nothing installed and takes about a second, so there is nothing to forget. The three big studies are slower, because each one rebuilds itself from the raw sales file: about half a minute for all three. Those run before each push, and again automatically whenever a study's own data changes.
  1. Rebuilds the analysis from the cleaned data. It does not read the published answer first, so it cannot be led toward agreeing.
  2. Compares its answer against the published data file. Catches the common failure: someone edits a script and forgets to re-run it, so the site keeps serving last month's answer.
  3. Reads the page and finds each figure by the sentence around it. Not by hunting for a bare number, which could match anything. It looks for the claim being made, then checks the number inside it.
  4. Fails if the sentence it expects has gone. If prose is reworded so an anchor no longer matches, that is an error, not a pass. A check that goes quiet when it stops understanding a page is worse than no check, because it reads as approval.
  5. Fails on any figure it does not recognise. Add a new number to a guarded page without declaring where it came from, and the script refuses it. That is how the guarantee stays complete over time.
  6. Confirms that retracted figures have not come back. Old numbers we have withdrawn are searched for across the site, and allowed only where the surrounding text is correcting the record.
  7. Exits with an error, which blocks publishing. This is a gate before a change ships, not a suggestion.
The home page is the hardest one to check. It quotes figures from seven separate studies, and its two charts are drawn by hand rather than from a data file, so nothing about it updates itself. Its script now checks all 45 of those figures, in the page and in both charts, and fails on any number that has no declared source at all. That last rule is what caught the mobile chart still showing county-wide figures the study had already withdrawn.

Pages without one
The other 21 pages. Their charts still come straight from the data. Their sentences are on their own.

This is not a lower standard of sourcing. Every figure on these pages was worked out from the same records and is cited in the page footer. The difference is what happens afterwards.

  1. A script or a document produces the figure once. Same data, same scripts, same care as anywhere else.
  2. Someone reads it and types it into the page. Checked by hand at the time, against the file it came from.
  3. The source goes in the footer. So a reader can go back to the original and work it out again.
  4. Nothing checks it again. Not when the data is refreshed. Not when the prose is edited. Not when a related figure elsewhere is corrected. The number stays until a person happens to look at it.

So the honest description of a figure on one of these pages is: correct when written, sourced, and unwatched since. That is a good deal weaker than the guarded pages, and you should read it that way.

Page Charts drawn from data Prose checked by a script
Checked in full
Homepage45 figures from seven studies, plus both hand-drawn chartsnoyes
Assessment regressivityyesyes
Reassessmentyesyes
Zoning, Explainedyesyes
IDA Overviewyesyes
IDA Projectsyesyes
IDA & Elmirayesyes
IDA-Owned Propertyyesyes
The IDA Boardnoyes
Charts live from data, prose unwatched
City Fiscal Healthyesno
Frozen Assessmentsyesno
City Budget Exploreryesno
County Budgetyesno
City–County Relationshipyesno
The Water Boardyesno
Multi-Year Trendsyesno
Fair-Share Mapyesno
Maintenance Rebateyesno
The Long Declineyesno
Tax Value per Acreyesno
All figures typed, nothing watching
PILOT Agreementsnono
Fiscal Decodernono
Why It Mattersnono
1940 Redlining Mapnono
Council Minutesnono
Contact Your Officialsnono
The Talknono
Retired pageRedirects to Tax Value per Acrenono
Glossarynono
Data & SourcesScanned for retracted figures onlynono

Every page, in full

The same panel that sits at the foot of each page, gathered here so you can read the lot without visiting them one at a time.

Each one says where that page's data starts, which script works out its figures, what gets filtered out on the way, and which choices were judgment rather than arithmetic. They carry rules and never figures — “drop any sale under $10,000” stays true when the data is refreshed, where a count would quietly go stale. Where a rule is enforced by a line of code, the recipe names that line, and a check fails the build if the code stops saying what the recipe claims.

Every source is downloadable twice over: the publisher's own copy, and a gzipped copy of the exact file this site used. Both matter, because the publisher's version moves on and one of them is a search form rather than a fixed file.

Loading the recipes…


One number, traced the whole way
A guarded figure, start to finish. The assessment page says cheap homes in Elmira pay about twice as much tax per dollar of what their home is worth as expensive ones. Here is where that came from.

Throughout this, the ratio means one thing: a home's assessed value divided by the price it actually sold for. A ratio of 0.60 means the assessor valued the house at 60% of what a buyer paid. The higher the ratio, the more tax you pay per dollar your home is really worth.

  1. Start with every recorded sale

    New York publishes the price of every property sale through its SalesWeb service. We keep a copy at data/raw/SaleswebExtract.csv, covering all of Chemung County.

  2. Narrow it to sales that can be compared

    Keep only the City of Elmira, only ordinary single-family houses, only arm's-length sales, only 2018 to 2025. Drop anything under $10,000 and any ratio above 5.0. Those are almost always data errors or transfers between relatives.

    What is left: 1,689 sales

  3. Work out each home's ratio

    Assessed value divided by sale price, using the assessment that was on the roll at the time of that sale, not a later one.

    The middle ratio of all 1,689: 0.631. The typical Elmira house was assessed at about 63% of what it sold for.

  4. Sort by price and compare the two ends

    Group the sales into price bands. Homes that sold under $50,000 had a middle ratio of 1.175. Homes at $150,000 and up had 0.457.

    Cheap homes were assessed above their sale price. Expensive ones at less than half. The raw gap: 2.573×

  5. Stop. That gap is too big, and part of it is our own doing

    Sorting by sale price puts sale price in two places at once: along the bottom of the chart, and underneath the ratio. A house that happened to sell cheap lands in the low band and gets a high ratio, for the same reason. The curve bends before any unfairness is involved.

    Sorting by assessed value instead does not fix this. It tilts the error the other way, because assessed value is the top of the ratio. So neither version can be measured against a flat line.

  6. Measure how much of the bend is our doing

    Find every house that sold twice within three years — 510 pairs. How far apart those two prices are tells you how much a single sale price strays from what a house is really worth.

    Then simulate a city with exactly that much randomness in its prices and no unfairness at all. Run the same method on it. It still produces a gap of 1.271×. That is the part we caused, now measured.

  7. Take it out

    2.573 ÷ 1.271 = 2.02×

    That is what the page publishes, and what it rounds to "about twice as much". The bigger, better-sounding 2.573× is not claimed anywhere, because part of it was never real.

  8. Write it out and draw it

    All of it goes into jcurve.json. The assessment page downloads that file and draws both curves, raw and corrected.

  9. Check it, including against our past selves

    audit_jcurve_figures.py redoes every step above from the raw sales file each time it runs. It also refuses to let a price band be drawn if too few sales fall into it, which is the mistake that produced the first version of this analysis.

    That first version pooled all eleven towns in the county together. Those towns assess on very different scales, so differences between towns were being read as unfairness within Elmira. It was wrong. We withdrew it, and the check now makes sure those old figures cannot quietly reappear anywhere on the site.


And one that no script can ever check
The PILOT page says Elmira College pays the city $5,000 a year, unchanged since 1994. That figure is as important as the 2.0×, and it is held up by something completely different.
  1. There is no dataset behind it

    The 2.0× came from 1,689 rows in a file. This came from an agreement signed in 1994 between the city and a college. There is no table to recompute, no roll to filter, nothing for a script to rebuild.

  2. What we have instead

    The City Chamberlain's office confirmed the amount and said it has not changed since the 1994 agreement.

    The payment appears in the 2026 adopted budget under account 412890, "Other General Department Income". Notably it is not booked with the PILOTs at all, which is itself part of the story.

  3. What we do not have

    The executed 1994 agreement itself. A records request for it is open, and the PILOT page says so in plain sight rather than implying we have read it.

  4. So checking it means something entirely different

    No script can confirm this figure, and adding one would not help. Verifying it means a person reading a budget PDF, or obtaining the agreement.

    If the city renegotiated the payment tomorrow, nothing in our pipeline would notice. The sentence would sit there being wrong until someone read the next budget.

  5. What holds it up, then

    Naming the source precisely enough that you can go and check it yourself, and saying out loud which part we have not seen. That is a weaker guarantee than the 2.0× has, and it should be read as weaker.

Both kinds of figure are on this site, side by side, and they do not look different. A reader cannot tell from the page which one is watched by a machine and which rests on a phone call to the Chamberlain's office. That is the gap the table above is meant to close, and it is the main reason this page exists.

Rules that exist because we got it wrong before
Each of these is a mistake that reached the live site, now written into the code so it cannot happen the same way twice.
Shipped three times

There are two Elmiras

The state files city parcels under the surrounding town. Ask the data for "Elmira" and you also get about 3,800 parcels that pay no city tax at all. On the 2025 roll that is 3,161 properties in the Town of Elmira, plus 625 utility assessments filed with the town. Including them makes the city look wealthier and less taxed than it is.

Every script now selects city parcels through one shared function. If it cannot find them it stops with an error, rather than quietly handing back an empty result.

Silent failure

The city's code is a number and a word

Elmira's official code is 070400. Read from a saved file it arrives as the number 70400. Read from the state's live service it arrives as the text "070400". Compare the wrong one and you match nothing at all — no error, just zero results.

One shared function now converts both forms before anything is compared.

Wrong for a year

Never merge the two tax rates

Elmira's rate can be given per $1,000 of assessed value or per $1,000 of market value. They are different numbers measuring different things.

Averaging them once put a miscalculated rate on this site for a year. Any page showing a rate now has to say which of the two it means, in text you can see.

Publishing limit

Nothing over 25 MB ships

The host refuses any single file above 25 MB, so the maps and data files are built to stay under it. That is why the site sends you small summary files rather than the whole roll.

The full roll is public. It is linked on the sources page, straight from the state.


What this page does not tell you
Knowing how a figure was built is not the same as knowing where the method is weak.

There are several places where a careful person could do this differently and reach a different answer. The state's published equalization rate for the whole roll says 56%, while the ratio measured from house sales lands anywhere from 46% to 63% depending on which sales and whose dollars. The 3D maps join parcel shapes recorded in 2021 to assessments from 2025. Sorting homes by sale price bends the assessment curve before any unfairness is involved.

Those have a page of their own: Known limits and judgment calls. Each entry says what we chose, why, what the choice costs, and what would change if you chose otherwise. It also lists the figures we have withdrawn, and what this site does not claim. If you are here because we asked you to check our work, start there.

Still to come: a figure-by-figure list covering the whole site, built from the checking scripts themselves.

Found something wrong? That is the point of publishing this. Tell us and we will look at it. Corrections get logged in public on the Data & Sources page, including the ones that embarrass us.