Skip to content
Methods

How the numbers are made

The definitions behind every report. They do not change between schools, trips or regions.

Waves

Before (T0)
Opens ten days before departure. The baseline.
Just after (T1)
Opens three days after return. The twenty-one items again, plus the experience module and one question on enjoyment.
Weeks after (T2)
Opens the number of weeks after return the school set for the trip, two to four as standard. The twenty-one items only. This is where any retention shows.
Six months after (T3)
Only on trips with the research module: opens twenty-six weeks after return, with the twenty-one items only. Reported beside the other follow-ups, always as change from before.

Counts

Invited
Students sent a link for a wave.
Responded
Students who finished the wave, with every core item answered. A student who answered some items and stopped is in no count and no figure.
Matched n
Students who finished both the baseline and the follow-up in question. Every change figure is computed on these students only, and the number is shown wherever a change is.
Suppression
Any group of fewer than ten students is withheld. For means and distributions this happens inside the database, which counts students per domain and per item, and the database also refuses to store a report that carries a figure for a smaller group, so a page cannot show one even by mistake.
Splits
Where students are split (by starting score, or by sex), a part is shown only when it has at least ten students and the students not shown number none or at least ten. If the rest would fall short, the smallest part shown is withheld as well, so a withheld part can never be worked out by subtracting the others from the whole.
Pooled
Across a school's trips, or an operator's, each student counts once: a pooled mean is the sum of each trip's mean times its students, divided by all the students, so a trip of ninety weighs nine times a trip of ten. Pooled outcome means are computed in the database over the trips' answers together, withheld per domain under ten. A trip joins a school's pooled change for a wave only when its own report shows every figure for that wave, so the pooled figure less the trips' own never leaves a withheld group to be worked out; a pooled group by sex is withheld when the students in it whose own trip withholds that group number between one and nine. Every pooled figure says how many students and trips it is over.
Same answer to every item
The share of students who gave one answer to all twenty-one core items at a wave. Shown under data quality; the answers are kept, not corrected.

Change

Points
Mean change on the four-point agreement scale between two waves, on matched students. A domain score is the mean of its items for that student; the overall score is the mean of every core item.
Against students not yet travelling
Where a comparison cohort exists, the reported figure is the difference in change: the travelling students' mean change minus the cohort's mean change over the same calendar window, before the cohort travels. The cohort answers on the trip's own dates at every wave, its baseline included; its own before-wave, just before it travels, is never used. A cohort linked after the trip's before-wave has closed has no baseline on those dates, so the trip is reported as change over time and the report says why. It is a difference between two changes, not a change, and every sentence says it as one. Students were not randomly assigned, so the two groups may differ in ways this does not remove.
95% interval
From the t distribution on the observed spread of change scores. Intervals built this way contain the true change in 95 of 100 repeats of the same measurement; any one interval either does or does not. Where it includes zero, this group is too small to tell the change apart from normal variation; that is not the same as no change.
No change worth reporting
Said only when the whole 95% interval for g lies within ±0.10, the smallest change treated as meaningful here. A point estimate near zero with a wider interval is reported as unresolved, never as no change.
Too small to matter
An interval that excludes zero but lies wholly within ±0.10: a change is there, and it is smaller than the smallest treated as meaningful. It is not counted among the clear changes a headline names.
g
The change divided by the standard deviation at baseline, with Hedges' small-sample correction. Where a comparison exists this is Morris's (2008) dppc2: the difference in mean change over the pooled baseline standard deviation of both groups. The standardiser is treated as fixed when the interval is computed, which makes the interval slightly narrower than it would be if its uncertainty were carried. It is reported so a change can be compared across domains and trips; the points figure is what it means in the scale's own units.
Smallest detectable effect
The smallest g this many matched students could detect at 80% power and 5% significance, from the observed change-score spread. An estimate inside it is below the design's resolution, not evidence of nothing. The sentence on how large a change a trip could pick up is read from this number.
Cliff's δ (with a comparison)
The probability that a travelling student's change exceeds a comparison student's, minus the reverse. Reported because the items are ordinal and this assumes nothing about their spacing; the words negligible, small, medium and large follow Romano's thresholds for this two-group form.
Up minus down (without a comparison)
The share of matched students whose score rose, less the share whose score fell. It is the one-group, paired form of the same idea and carries no size word, because Romano's thresholds were set for two groups.
Did it last
The change at follow-up read against the change just after: at least 0.9 of it is all, at least half is most, anything above zero is some, and zero or below is none. The two figures are on the students who answered each wave, so the groups overlap but are not identical, and when the follow-up interval includes zero no claim about lasting is made. Across a school, both figures are over the same trips, those with a follow-up, and the weeks are named only when those trips share them.
Unit of analysis
Within a trip, the student. Across a school's trips, students on the same trip share a leader, a group and a week, so they are not independent: the school interval is widened by the design effect 1 + (m − 1)ρ, where m is the mean number of students per trip and ρ the share of the variation in change that lies between trips, and takes its t value on the number of trips less one, since the trips are the independent units: a school of two or three trips can rarely tell a change apart from the difference between its trips. Between trips, for the design question below, the unit is the trip.

Students one at a time

Higher, same, lower
The share of matched students whose score rose, stayed exactly the same, or fell, counting any movement. Only the next measure says a student improved.
Beyond measurement error
Movement larger than the scale's own noise, after Jacobson and Truax: the standard error of measurement is the baseline standard deviation × √(1 − α), the standard error of a difference is √2 times that, and a change beyond 1.96 of it is reliable at the 5% level. On a short scale the threshold is large, so this share is always smaller than the raw one.
α
Cronbach's alpha at baseline, the internal consistency of a domain's items. Short scales (four to six items) cap α by their length; the item count is shown beside it.
Floor and ceiling
The share of matched students at the bottom or top of the scale at baseline. A student at the top cannot register a gain, a student at the bottom cannot register a decline.
Change by sex
Where the register records sex, the overall change for each group, under the split rule above. Students with no entry are in no group but count in what is left over. Sex is the only identifying field collected; it is optional on the register.
Change by starting level
Matched students cut into three bands on their baseline score, at the scores that would split them into thirds, under the split rule above. Students on a cut's score all go in the band above it, so where many share a score the bands are unequal or one is empty; each band is named by the scores it covers, and an empty one says none. Low starters rise on remeasurement and high starters fall whatever happens (regression to the mean), so where a comparison cohort exists the same cuts are applied to it and both are shown; a real difference is a gap between the two, not a large number in one. Across a school's trips, each trip's students are banded at that trip's own cuts and the bands added up, counting only trips whose own report shows every band, so a band one trip withholds is never the school's band less another's.

Design and outcome

Enjoyment
One question after the trip, "Overall, I enjoyed the trip.", reported as the share who agreed. It sits beside experience quality so the two can be seen to differ, and it is part of no score.
Experience quality
Post-trip ratings of preparation, novelty, challenge, autonomy, educator support, group functioning, reflection and transitions. Answered once, days after the trip. It is a rating of how the trip was designed and delivered and is never an impact score.
Priorities
Fixed rules, one per dimension depending on what exists. Where the platform median exists for a dimension (once three or more operators have measured trips), the dimension is a priority when it is 0.11 or more behind that median, and a strength when 0.11 or more ahead of it. Where there is no median, it is a priority below 2.88 and a strength at 3.25 or above. Priorities are ranked by gap, then by score, and capped at 3. No text is generated by a model.
Did the design predict the outcomes
Five pairings fixed before any data existed, reported only once 30 trips across 5 operators have a complete post-trip wave, because the unit of analysis is the trip. Each trip contributes the mean change of its matched students from before to just after on each domain: the same quantity whether or not it had a comparison cohort, so a controlled trip's difference in change is never mixed in. The pairings:
  • Challenge and Perspective and accomplishment: Trips whose students rate the activities as better pitched show larger mean gains in perspective and accomplishment.
  • Autonomy and Agency and self-management: Trips whose students report more real decisions show larger mean gains in agency and self-management.
  • Group and Social connection and belonging: Trips whose students report groups where everyone took part show larger mean gains in social connection and belonging.
  • Reflection and Learning and educational value: Trips whose students report more time to reflect show larger mean gains in learning and educational value.
  • Preparation and Agency and self-management: Trips whose students felt less prepared for what was coming show smaller mean gains, or falls, in agency and self-management.

What every figure is

What students say about themselves, on the same items at each point. A trip can change what a student counts as coping well, which can make a real gain look flat; nothing here corrects for that. The instrument.