A new paper audits the Capability Sprints' compute-hours-per-output metric and finds a facilitator leaderboard's accidental staggered rollout moved usage, not the analyses behind it.
A rollout accident let the Adaptive Metrics Lab isolate visibility from measurement, and only visibility moved compute-hours across the Capability Sprints
The Adaptive Metrics Lab renewed its shared GPU cluster for a second year on the strength of one figure: compute-hours-per-output, the median compute logged against every Capability Sprints capstone submission. A new paper from the School of Continuous Improvement finds the figure rose sharply only once it became something Sprint facilitators could actually see.
Dr Bram Ntuli and Dr Mirela Hanke traced the pattern to an uneven leaderboard rollout across the program’s five short courses: three gained a facilitator-visible ranking a full teaching period before the remaining two. Reading submission logs across the eight weeks either side of that launch, they compared how much compute each capstone used against whether its result was new or a rerun of a participant’s own earlier work. Compute-hours climbed roughly 3.5-fold in the three Sprints that could see the ranking; the two still working blind held steady. The share of submissions producing anything distinguishable from a participant’s prior work barely moved either way.
The School reads the result as exactly the kind of check a self-reported efficiency figure ought to invite before its next renewal cycle.
Once a number is visible, people respond to the number. We just wanted to know what they were responding to.
— Dr Ntuli, Lecturer and Convenor of Improvement Grand Rounds
“A metric that only reports usage was always going to need a companion,” said Associate Professor Casimir Beng, Lead of the Adaptive Metrics Lab. “The Lab intends to treat the novelty check as exactly that.”
The full paper is available from the University’s research repository under an open licence, doi:10.5555/slop.yr93gz.
