When the questions change, can we still trust our analysis?

Published on August 17, 2026

How and when to compare apples to oranges.

Here’s a common scenario when working in new areas of system change: you develop a custom survey, administer it and realise the initial questions do not quite capture the problem. So, you tweak them.

But this raises a crucial question for your organisation's reporting: Can we compare the results from surveys over time if the wording shifts?

Take a standard digital access survey, for example:

Old question: "Do you have the internet at home?"

New question: "Do you currently have internet access? If yes, how do you usually access it?"

The new question allows for a greater variety of living circumstances, but can we still utilise the old data now that the question has changed? Here is how to manage these shifts without losing the value of your historical data.


Did the fruit get heavier… or did you change the scales?

Changes in survey questions, whether superficial tweaks to words or substantive alterations, directly affect data comparability. Shifts in response options or the way surveys are administered influence respondent behaviour and can introduce bias.

Recognising these changes is essential to maintaining the validity of your longitudinal insights – ensuring the trends you see reflect real-world changes, not just a shift in how you asked the question.

Common survey changes

Understanding exactly how your survey has evolved is the first step to fixing the data. Question evolution usually falls into one of these categories:

Change Type What It Involves Example
Wording Adjusting phrasing for clarity 'Experienced housing insecurity?' vs. 'Difficulty maintaining stable housing?'
Scale Changing the number of response options 3-point engagement scale vs. 5-point Likert scale
Response Options Adding or removing categories Including a 'Prefer not to say' option
Mode and Sampling Changing how the survey is administered Moving from staff-led interviews to online forms

Your step-by-step action plan

Managing these changes requires clear documentation. Follow these steps to ensure your data remains robust.

  1. Audit all changes
  2. Transform the data
  3. Test the results
  4. Document and visualise

1. Audit all changes

Document versions, dates and specific wording shifts in a dedicated change log.

Before you attempt any harmonisation method (explained in further detail below), there is one inexpensive, high-value practice that makes everything else possible and that almost every organisation skips.

Treat the version of your questionnaire as data and attach it to every single response you collect.

This is what we call iteration metadata.

Iteration metadata is a small set of fields that travels with each record, recording which version of the survey produced it. It is the hinge between recognising that a question changed and being able to do anything useful about it. You cannot harmonise what you have not tracked.

Why it matters

Questions can change without the data team knowing. A program officer rewords something in the form builder, a volunteer tweaks the phrasing during a busy intake period and it may only surface months later when the numbers look odd. This risk is amplified in environments with frequent staff turnover, where institutional knowledge of why a question was worded a particular way leaves with the person.

Iteration metadata makes the discontinuity visible. Once every record carries its version, you can separate the two populations, see exactly where they join and make a deliberate decision about whether and how to bridge them.

It is also the only part of this you genuinely cannot fix retrospectively. You can rescale and recode historical data after the fact. You usually cannot reconstruct which version a three-year-old response came from once that information is gone. Capturing it at the point of collection costs almost nothing; recovering it later is often impossible.

How to use iteration metadata once you have it

The metadata is only worth capturing because of what it lets you do.

  • Filter to a single version for a clean within-version trend, with no contamination from the rewording.
  • Segment charts by version and annotate the join with a clearly labelled break line, so partners and funders see the change rather than reading straight through it.
  • Join the change log to your analysis to pick the right harmonisation method for each segment, rescaling in one place, recoding in another.
  • Run version against your outcomes as a sense-check. If an "improvement" lines up exactly with the date a version changed, you are most likely looking at an artefact of the question, not a result in the world.

2. Transform the data

Choose the appropriate continuity or rescaling technique for your specific change type.

Not every change needs every method. Which one you use comes down to the question sitting under this whole guide: is the new version a better measurement of the same thing, or a changed idea of the thing? Match your change to the approach below, then carry out only the steps it points you to.

Your change The method to reach for
Wording only (same concept) Usually no transform. Confirm the change was cosmetic with a trend-break check. If the meaning shifted, treat it as a scale or response-option change below.
Scale (e.g. 3-point to 5-point) Rescale the old scores onto the new scale. Use equipercentile linking if the scale points are not evenly spaced.
Response options added or removed Recode responses into the categories both versions share.
Mode or sampling (who answered changed) Weight the sample so the mix of respondents stays comparable over time.

Whichever you pick, record the decision in your change log from Step 1, so the reasoning travels with the data.

What is equipercentile linking?

Equipercentile linking is a statistical way to compare scores from two different tests or surveys by matching their percentile rankings.

If a score on Survey A and a score on Survey B both put a respondent in the exact same percentile (for example, the top 10%), those scores are considered equivalent – even if the raw numbers or scoring scales are completely different.

Harmonise the variables

Transform variables to a common format by normalising scales or recoding categories.

Harmonising means transforming the data so old and new responses can sit in one series. For the changes above, that is almost always one of two moves.

  • Rescale – when the scale changed
    Map the old scores onto the new scale, keeping the ends fixed so the midpoint of one lines up with the midpoint of the other. When the scale points are not evenly spaced, use equipercentile linking instead of simple arithmetic.
  • Recode – when the response options changed
    Collapse responses into the categories both versions share.

Always transform a copy, never your raw data, and note which records you changed. If you ran both versions side by side for a period, use that overlap to check the transform reproduces the original pattern before you trust it.

Apply weighting adjustment

Account for sampling or mode changes to ensure demographic comparability.

Weighting matters when what changed was who answered, not what you asked. In the digital access survey, an online form is easiest for people who are already well connected, so reported access rises, not because the community changed, but because the sample did.

By weighting, we correct this by giving each group its proper share. If younger respondents are over-represented online, their answers are counted for proportionally less, so the sample matches a consistent profile across time.

Also required is something to weight towards, for example: known proportions for age, location or another characteristic, drawn from your intake records or ABS data. If you do not have reliable figures, do not invent them. The honest alternative is to flag the mode change as a limitation on the chart and in your notes, so readers can weigh the trend accordingly.


3. Test the results

Use statistical tests to confirm if data changes are real-world trends or simply the result of a new online format.

Once the data is harmonised, one question remains: is the change you see real, or an echo of the change you made to the survey? Testing for a trend break looks for a jump that lines up with a version change rather than with anything happening in the world.

Plot the series with the version change marked. If the line moves exactly at that point, and nowhere else, be suspicious; that’s the signature of an artefact, not a trend.

To confirm it, compare responses from an overlap period, or use a statistical test: a chi-square test for categories, a t-test for averages, or interrupted time-series analysis for longer runs. Each tells you whether the difference is larger than normal variation. If you do not have these skills in-house, this is a good moment to draw on the support options at the end of this guide.

Conduct sensitivity analyses

Evaluate the robustness of your results by systematically testing your harmonisation assumptions.

Harmonising involves judgement calls – where you drew category boundaries, how you rescaled, which records you weighted. A sensitivity analysis asks a simple question: if you had made a different reasonable choice, would your conclusion change?

Redo the analysis with a plausible alternative. Recode the categories a slightly different way, or rescale on a different assumption, and compare the result.

If your headline finding holds across these variations, you can report it with confidence. If it flips depending on how you harmonised, you do not have a finding – you have an artefact of your method, and should say so rather than choose the version you prefer.


4. Document and visualise

Record all decisions and clearly annotate your charts to enhance transparency for partners and funders.

Your decisions are only trustworthy to others if they can see them, and you have already done most of the work: the change log from Step 1 is your record. You do not need a second one.

On any chart that crosses a version change, mark the join. A clearly labelled break line – 'question changed here' – lets partners and funders read the trend with the shift in mind, rather than reading straight through it as though nothing happened.

Note the method you used, and any limitation, in a short caption or footnote: a weighting you could not fully verify, or a wording change you judged cosmetic. Being open about a break is far more credible than a smooth line that quietly hides one.


So, can you compare apples to oranges?

The organisations that can answer this question are the ones who saw it coming and stamped their data accordingly. By understanding the nature of your changes and taking calculated steps (like running concurrent surveys or statistically adjusting scales) you can successfully harmonise your data and continue delivering impactful, evidence-based insights.

The quickest win here is also the most powerful. If you change nothing else, start recording which version of the question produced the answer directly in your spreadsheet.

Looking for extra support for your own data projects?

  • Apply for a DCN data project if you are a small-to-medium not-for-profit.
  • Watch the webinar on the ABS Lifecourse project for details on accessing data analysis support from the Australian Bureau of Statistics.
  • Consider engaging a university partner for longer-term projects requiring high-level statistical capabilities.