Review Reports
- Aldo Seffrin 1,
- Pantelis Theodoros Nikolaidis 2 and
- Beat Knechtle 4,5,*
- et al.
Reviewer 1: Anonymous Reviewer 2: KrisztiƔn Havanecz
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsReview report
I would like to thank the authors for providing me with the opportunity to review this study. This study presents a comprehensive and valuable dataset containing more than 2.7 million athlete-race records from IRONMAN® and IRONMAN® 70.3 races held between 2002 and 2026. In particular, the inclusion of swimming, cycling, and running times—as well as T1 and T2 transition times—as separate variables is a significant contribution that sets this dataset apart from similar studies in the literature. In addition, the simultaneous integration of multiple methods in the data validation process enhances the reliability of the dataset.
However, I believe it would be beneficial to consider the minor revision suggestions listed below and revise the manuscript in order to improve the academic quality of the article and make it more effective.
- Creating the abstract section
- The first sentence is quite long and conveys two different messages (gaps in the literature + limitations of existing datasets). This makes it difficult for the reader to grasp the main objective on the first reading. Therefore, the first sentence should be split into two shorter sentences so that the gaps in the literature are addressed first, followed by a clear statement of the study’s objective.
- The purpose of the study must be stated in separate and clear sentences.
- Currently, the validation results are presented one after another with three different percentage figures. Instead, the summary could begin with a general conclusion: “The dataset demonstrated excellent internal consistency and inter-source agreement.” The most important validation findings could then be listed. This will help the summary flow more smoothly.
- Creating the introduction section
- The final paragraph provides a broad list of potential applications for the dataset. However, some examples—such as examining post-COVID-19 return-to-competition patterns or the impact of shoe technology on running performance—give the impression that the dataset contains variables that would directly support these analyses. These examples should be phrased more cautiously. Furthermore, since the paragraph is structured as a list of potential research topics rather than highlighting the dataset’s core contributions, it would be beneficial to reorganize it in a more focused and coherent manner.
- Although the necessity of the study is successfully explained in the introduction, the purpose of the research and, if applicable, the research hypothesis are not clearly stated. Even though this is a data descriptor article, specifying a clear research objective and/or research question at the end of the introduction would strengthen the focus of the study and help the reader better understand the article’s main contribution.
- Insufficient detail in the material method section.
- It is stated that the dataset was compiled from an independent data source using the official IRONMAN® results platform. However, if records for the same athlete or the same race appear in both sources, a more detailed explanation is needed regarding how duplicate records are identified and the method used to remove them.
- The rationale for combining IRONMAN® and IRONMAN® 70.3 races into the same database should be explained more clearly. This is because these two formats involve different physiological demands and strategies. Therefore, it would be helpful to provide recommendations regarding distance-based analyses for researchers who will use this dataset.
- The dataset covers a fairly long period of time. This is a significant strength of the study. However, it should be clarified whether any changes were made to the age categories of the competitors during those years.
- Strengthening the Limitations Section
- It should be clearly stated that the dataset does not include environmental variables—such as course profile, elevation, air and water temperature, wind conditions, and drafting rules—that can significantly affect triathlon performance, and the limitations this may pose for future secondary analyses should be discussed.
- Creating the conclusions section
- The Conclusions section successfully summarizes the size of the dataset and its potential scientific contribution. However, it should be made clearer on what criteria strong statements such as “the largest and most comprehensive dataset in the field” are based, or such statements should be phrased more cautiously.
General Assessment
This study presents a high-quality and comprehensive dataset that fills an important gap in triathlon literature. The data collection and validation process appears to have been successful, and the suggested revisions are primarily aimed at strengthening the explanations regarding the methodology and the use of the dataset. Therefore, I believe the article is suitable for publication following minor revisions.
Recommendation: Minor revision.
Author Response
Response to Reviewers — IRONMAN® Data Paper R1
Manuscript: “A Validated, Population-Scale Dataset of 2.7 Million IRONMAN® Triathlon Records with Separated Transition Times (2002–2026)” Manuscript ID: sci-4491984 Journal: Sci (MDPI) R1 decision received: 2026-08-19 R1 response date: [TO BE FILLED at resubmission]
Note on Open Review: this manuscript is under Open Review, so this response is published alongside the paper. It is written to be read by anyone, not only by the reviewers.
We thank the Editor and both Reviewers for reports that were specific enough to be acted on rather than merely acknowledged. Three of Reviewer 2’s comments asserted that we had made an error. We checked all three against the data. Two were correct, and checking the third — which does not reproduce — led us to a defect neither reviewer had seen. Working through the remaining comments surfaced two further errors that no reviewer raised. All four are reported below.
We have accepted every comment. There is no point on which we argue with a reviewer’s recommendation.
The substantive changes are:
- Cross-source validation expanded from 6 races to the entire overlap between the two sources — 559 race-years spanning 2003 to 2026, 340 full-distance and 219 half-distance, yielding 4,732,776 matched athlete-discipline pairs. This addresses Rev2-Q1 at its root: with the whole overlap compared, there is no sample to justify and no representativeness to argue.
- The independence of the second source is conceded (Rev2-Q13), and what the agreement statistic does and does not establish is now stated wherever it is reported.
- Eleven new analyses, among them T2 missingness mirroring T1, age-category stability across the covered years, half-distance plausibility bounds, merge-rule and event-coverage counts, the condition of the country field, and a three-rule athlete-matching comparison that revises an explanation we had given.
- Five figures regenerated and one added; Table 1 and Table 2 rebuilt from generated output rather than hand-maintained.
- A limitations section restructured around the five points Rev2-Q25
A note on how the numbers in this letter were produced. Every quantitative claim in the revised manuscript is now checked against a persisted result file by a script that runs after any regeneration and exits non-zero on a mismatch; it currently verifies 100 claims. We built it during this revision because two of the errors we found were caused by figures that had been correct when first computed and were never recomputed afterwards.
Editor’s Letter
R1-Ed-1 — Self-citation rate
“During the technical check of your manuscript, we noticed that a high proportion of the cited references belong to you or your co-authors Refs. 2,3,4,5,7,12,16, which is a self-citation rate of about 43.75%. Could you please check whether the inclusion of each of these references is appropriate?”
Response. We checked, and the check produced an answer we did not expect: the rate was higher than the office reported.
Reference 15 (Loosli et al., Frontiers in Sports and Active Living, 2025) is also a self-citation — Nikolaidis, Andrade, Rosemann and Knechtle are among its ten authors — and was not flagged. The reason it was not flagged is our own error: the bibliography entry truncated the author list, so the reference rendered under the first author alone and the four co-authors were invisible to anyone counting. The submitted manuscript’s true self-citation rate was therefore 8 of 16 (50%), not 43.75%. We have corrected the entry, which also corrects the first author’s forename.
We then mapped each flagged reference to the claim it supports and classified it as load-bearing or not. Reference numbers below are those of the submitted version, which is what the office counted; restructuring the Introduction has since changed the order of first appearance and therefore the numbering.
Removed — three references that carried no argument:
|
Ref |
Reason |
|
4 — Knechtle et al. 2025 (environment) |
Supported the same Introduction sentence as reference 3, on the same point. One citation suffices. |
|
12 — Knechtle et al. 2024 (sex differences) |
A single mention inside an enumeration of possible research directions. The enumeration lists topics; it does not argue from sources. |
|
16 — Sousa et al. 2021 |
The same, in the same enumeration. |
Retained — each with a claim that fails without it:
|
Ref |
The claim it carries |
|
2 — Nikolaidis et al. 2023 |
The comparator for the dataset’s size. The statement that this resource is 3.3 times larger than the largest previously published IRONMAN® analysis requires the analysis it is measured against (823,459 records). |
|
3 — Knechtle et al. 2025 (origin) |
The scale of prior full-distance work (677,320 records), which establishes what “population-scale” meant before this dataset. |
|
5 — Knechtle et al. 2019 |
The only prior population-scale measurement of transition time we are aware of — that the fastest finishers spend 0.9% of race time in transitions against 2.2% for the slowest. It is the quantitative basis for treating transitions as worth separating. |
|
7 — Rüst et al. 2014 |
The study this dataset exists to answer. It is the one dedicated analysis of transition times at IRONMAN® distance we are aware of, it was restricted to the annual top ten World Championship finishers, and it could not separate T1 from T2. Removing it would remove the paper’s motivation. |
|
15 — Loosli et al. 2025 |
The review behind the statement that the research directions we list are recognised gaps rather than our own suggestions. |
We have simultaneously added four external references, in answering Rev2-Q4 and Rev2-Q26: Laursen & Rhodes (2001), Laursen et al. (2006), Millet & Lepers (2004) and Jeukendrup (2011).
The revised manuscript cites 17 references, of which 5 are self-citations — 29.4%, against a true submitted rate of 50%.
Reviewer 1
We thank Reviewer 1 for a reading that identified where the methods section was thin. Most of the ten points asked us to state something we knew but had not written down, and in two cases (Q6, Q8) writing it down required computing it, which we had not done.
R1-Rev1-Q1 — Abstract opening sentence carries two messages
“The first sentence is quite long and conveys two different messages […] the first sentence should be split into two shorter sentences so that the gaps in the literature are addressed first, followed by a clear statement of the study’s objective.”
Response. Done, in the order the Reviewer specifies. The Abstract now opens with the field’s use of long-distance triathlon as a model, states the constraint that existing datasets impose as a separate sentence, and follows with the aim. Splitting the opening sentence and adding an explicit aim did not lengthen the Abstract: it is 205 words against 207 in the submitted version.
R1-Rev1-Q2 — State the purpose in separate, clear sentences
“The purpose of the study must be stated in separate and clear sentences.”
Response. An explicit aim sentence has been added as the third sentence of the Abstract: “Our aim was to assemble, validate, and describe an openly available dataset that removes both constraints.” The wording was chosen after settling the position on “validated” that Rev2-Q1 raises, so that the two are consistent.
R1-Rev1-Q3 — Validation results as three consecutive percentages
“Instead, the summary could begin with a general conclusion […] The most important validation findings could then be listed.”
Response. Adopted. The Abstract now leads with the conclusion — “The dataset showed close correspondence between sources and high internal consistency” — and the figures follow it. The figures themselves changed, because the validation was rerun over a far larger comparison (Rev2-Q1) and because one of them was a rounding that read as a contradiction (Rev2-Q19); the Abstract carries the corrected values.
R1-Rev1-Q4 — Introduction final paragraph overpromises applications
“Some examples—such as examining post-COVID-19 return-to-competition patterns or the impact of shoe technology on running performance—give the impression that the dataset contains variables that would directly support these analyses.”
Response. The Reviewer is right, and we have taken the sharper of the two available routes. Rather than hedging each example, the research directions are now explicitly divided into those the schema supports and those it does not:
Others would require linkage to information the dataset does not hold: return-to-competition patterns after the COVID-19 disruption can be described in aggregate but not attributed, since no field records why an athlete was absent, and any study of advanced footwear technology would need external data on what athletes wore, which no results platform publishes. We list the second group to mark the boundary rather than to claim it.
The paragraph also moved. It had been duplicated between the Introduction and the Discussion; it now appears once, in Discussion §Research directions, which is what Rev2-Q7 asks for from the other direction.
R1-Rev1-Q5 — No clear research objective at the end of the Introduction
“[…] specifying a clear research objective and/or research question at the end of the introduction would strengthen the focus of the study.”
Response. An aim sentence was present in the submitted Introduction, but both Reviewers missed it — which told us the problem was placement rather than absence. It sat before a block of methodological description that Rev2-Q7 asks to be moved out. Moving that block leaves the aim as the closing statement of the Introduction, which is where both Reviewers looked for it.
R1-Rev1-Q6 — How duplicates across sources are identified and removed
“[…] if records for the same athlete or the same race appear in both sources, a more detailed explanation is needed regarding how duplicate records are identified and the method used to remove them.”
Response. The rule is now written out in Methods §Data merging, with the counts it produces:
- Deduplication is at the race-year level, not the athlete level. Overlap is detected by normalising event names — stripping year, brand prefixes and championship designations — and matching on normalised name, year and race type.
- When a race-year is present in both sources, the official record is retained and the supplementary record discarded, without exception.
- Applied across the two sources, 559 race-years were present in both and resolved in favour of the official source in every case, against 613 present only in the official source and 382 present only in the supplementary source — 1,554 distinct race-years in the merged dataset.
We also report what the normalisation contributes, because it is the step the rule depends on: matching raw event names identifies none of the 559 collisions, the generic rules identify 537, and the championship-prefix rules account for the remaining 22.
R1-Rev1-Q7 — Rationale for combining IRONMAN® and IRONMAN® 70.3
“The rationale for combining IRONMAN® and IRONMAN® 70.3 races into the same database should be explained more clearly […] it would be helpful to provide recommendations regarding distance-based analyses.”
Response. Both parts are now addressed, in the two places they belong.
The rationale is in Methods §Data consolidation: a shared schema is precisely what makes direct comparison across distances possible, and distance-specific datasets preclude it by construction. The race_type field preserves the distinction inside the file, so pooling remains a decision the analyst takes rather than one the data imposes.
The recommendation is in Discussion §Practical guidance for reuse, and we have given it the quantitative grounds the Reviewer’s own argument implies. The median finish among finishers is 5:53:16 at the half distance against 12:25:37 at the full distance, and the determinants of endurance performance are not the same across a range that wide (Laursen & Rhodes, 2001). Analyses should stratify by race_type or justify pooling explicitly.
R1-Rev1-Q8 — Did age categories change over the covered years?
“[…] it should be clarified whether any changes were made to the age categories of the competitors during those years.”
Response. We did not know, so we checked rather than asserting stability. Distinct age-group values were tabulated by race year across the whole dataset.
No boundary moves at any point. The five-year bands from 18-24 to 75-79 are present in every year from 2002 to 2026. What varies is only which of the oldest bands is populated: 80-84 is absent from 2006 to 2008, and 85-89 appears from 2015 to 2018 and again from 2022 onward. This reflects whether any athlete in those bands finished a given season, not a change in the classification. The finding is reported in Methods §Data consolidation.
R1-Rev1-Q9 — Environmental variables absent from the dataset
“It should be clearly stated that the dataset does not include environmental variables—such as course profile, elevation, air and water temperature, wind conditions, and drafting rules […]”
Response. Stated in Limitations, in the Reviewer’s own terms:
The dataset holds no course profile, elevation, air or water temperature, wind, or drafting regulation, so analyses of performance across events cannot control for the conditions under which those performances were produced.
We have added the temporal counterpart, which Reviewer 2 raises as the fifth item of Rev2-Q25: courses, qualification rules and participant populations changed over the 24 years covered, which makes any longitudinal comparison a comparison between eras as much as between athletes.
R1-Rev1-Q10 — Justify or soften “largest and most comprehensive”
“[…] it should be made clearer on what criteria strong statements such as ‘the largest and most comprehensive dataset in the field’ are based, or such statements should be phrased more cautiously.”
Response. We have separated the two claims, because they are not equally supportable.
“Largest” is retained with its comparator stated. At 2,706,922 records the dataset is more than three times the size of the largest previously published IRONMAN® analysis (823,459 records, Nikolaidis et al. 2023), and it spans both race distances rather than one. That is a checkable statement and it now appears with the number it is measured against, in the Discussion and in the Conclusions.
“Most comprehensive” and “most complete” are removed. They are the weaker claims, they are the ones Rev2-Q8 attacks from the coverage side, and we cannot demonstrate them — as the answer to Rev2-Q8 explains, no register exists against which event-level completeness could be computed. The Conclusions now name the resource’s weaknesses alongside its strengths.
Reviewer 2
We thank Reviewer 2 for a report that identified the manuscript’s central weakness correctly. The opening observation — that the reporting issues “limit the interpretation of the dataset as fully ‘validated’ and ‘population-scale’” — is one argument arriving from several directions, and we have answered it as one argument rather than hedging each mention separately.
Our position, stated once and referenced below. “Validated” is retained, and earned rather than defended: the cross-source comparison has been rerun over the entire overlap between the two sources instead of a six-race sample. “Population-scale” is retained with its meaning narrowed to what we can support — the scale of records assembled — and the event-level completeness the term might imply is declared unquantifiable, which is the route the Reviewer explicitly offered.
R1-Rev2-Q1 — “Validated” too strong for 6 races from one season
“[…] the term ‘validated’ appears relatively strong considering that cross-source validation was performed on 6 races from one season. The authors should provide stronger justification for this terminology or consider more cautious wording […]”
Response. We chose the first of the two routes the Reviewer offers, and provided the justification by enlarging the evidence rather than by arguing about the existing evidence.
The submitted validation used 6 races, all from the 2024 season. A census of the two raw sources showed that 559 race-years are present in both, spanning 2003 to 2026, 340 full-distance and 219 half-distance, with roughly one million records available for pairing on each side. The comparison has been rerun over all of them.
Results, now in Results §Technical validation and Table 2:
- Record counts from the two sources differ by two or fewer athletes for 484 of the 559 race-years (86.6%).
- Athlete matching yields a median rate of 3% per race-year under case-insensitive exact matching and 96.4% under an order-invariant rule.
- Among matched athletes, 4,732,776 athlete-discipline pairs were compared across all six disciplines. 76% agree exactly, to the second, ranging from 98.50% for the run split to 99.10% for the swim.
- Agreement is not uniform over time, and we report this rather than pooling it away: the median race-year agrees exactly for 100% of pairs and the fifth percentile for 99.21%, but 2017 sits at 86.3% and 2014 at 95.1%, against 99.9% or above from 2020 onward.
One caution, which we state in the Methods and repeat here. The 98.76% figure is not a like-for-like revision of the 99.7–99.9% reported in the submitted version. It covers 559 race-years instead of 6, and it is computed on an order-invariant match, which is a more permissive rule admitting pairs the earlier exact match never saw — including the 2017–2019 seasons, which matched at essentially zero before. A reader comparing 99.8% to 98.8% would conclude the data had got worse. The measurement got broader.
R1-Rev2-Q2 — “Five continents” may be four
“This statement that the 6 validation races represented five continents should also be checked. Based on the events listed in the methods, only 4 continents appear to be represented.”
Response. The Reviewer is correct. The six races were Florida, Frankfurt, Brazil, South Africa, 70.3 Oceanside and 70.3 Louisville — four distinct continents, with three of the six in North America. The claim appeared in four places: the Abstract, Methods, Results and the Table 2 caption.
We have not corrected the number, because the sentences containing it no longer exist. The expanded comparison under Q1 replaced them, and the validation set is now described by what it is — every race-year both sources hold, across 24 seasons and both distances — rather than by a continent count.
R1-Rev2-Q3 — Add IRONMAN to the keywords
“Please consider adding IRONMAN to the keywords”
Response. Added. The keyword list now reads: triathlon; IRONMAN®; open dataset; transition times; data validation; endurance performance; race results; reproducibility. The registered trademark symbol is carried as it is on every other mention in the manuscript.
R1-Rev2-Q4 — “Extreme physiological demands” not described
“[…] the manuscript refers to the ‘extreme physiological demands’ of IRONMAN® triathlon, but these demands are not further described. […] This does not need to become an extensive physiological review; one or two well-supported sentences would be sufficient.”
Response. Two sentences added to the Introduction, and no more, per the Reviewer’s own constraint:
Those demands are of a kind that the usual determinants of endurance performance do not fully capture: beyond roughly four hours of competition, maximal oxygen uptake and the anaerobic threshold cease to predict performance well, and fuel and fluid provision, substrate availability, and electrolyte balance become limiting in their own right [Laursen & Rhodes 2001; Jeukendrup 2011]. Competition over eight to seventeen hours also imposes a sustained thermoregulatory load [Laursen et al. 2006] and a progressive loss of neuromuscular function across the three disciplines [Millet & Lepers 2004], so that the athlete who begins the run is not physiologically the athlete who began the swim.
All four references are external to the author list, which also serves the Editor’s query.
R1-Rev2-Q5 — Rationale for separating T1 from T2
“[…] the importance of separating T1 and T2 should be explained more clearly. […] T1 occurs after swimming and before cycling, whereas T2 follows prolonged cycling before running.”
Response. The asymmetry the Reviewer describes is now stated in the Introduction as the reason the separation matters:
T1 follows the swim and precedes the cycle, so it is dominated by the change of equipment and by the shift from horizontal to upright posture. T2 follows several hours of cycling and precedes the run, placing it at the point where the run-off-bike penalty originates; it therefore reflects accumulated fatigue as much as logistical efficiency. Treating them as a single combined interval discards that distinction.
The point is not only conceptual, and the dataset now demonstrates it. Among athletes who did not finish, T1 is present for 75.1% of records but T2 for only 42.8% — because an athlete who abandons on the bike leg has passed through T1 and never reaches T2. A combined interval cannot carry that information.
R1-Rev2-Q6 — “Only” and “first” used cautiously
“[…] strong novelty statements such as ‘only’ and ‘first’ should be used cautiously unless supported by a comprehensive search.”
Response. We considered documenting a search to support the claims, and did not, because no systematic search was performed and describing one after the fact would not be honest. The claims are softened instead, to exactly what we can stand behind:
- “The one dedicated analysis of transition times at IRONMAN® distance that we are aware of […]”
- “We are not aware of a peer-reviewed, openly available, and formally validated dataset for IRONMAN® triathlon […]”
We have kept the corresponding claim in the Conclusions in the same register, and the one claim that is checkable — that this is the largest openly available resource of its kind — is stated with its comparator rather than hedged.
R1-Rev2-Q7 — Methodological content in the Introduction
“[…] the final part contains substantial methodological information. Most of this would fit better in the materials and methods section. The introduction should preferably conclude with a clear knowledge gap and concise study aims.”
Response. Done, and the block turned out to need deleting rather than moving. Its content — the dataset’s composition and a list of potential applications — duplicated Results §Dataset composition and Discussion §Research directions almost sentence for sentence. Relocating it to the Methods would have created a third copy. The Introduction now closes on the knowledge gap and the aim.
This single edit also answers Rev1-Q4 and Rev1-Q5.
Separately, the Methods gained an explicit statement of the quality-assessment structure, which the submitted version left implicit: four dimensions, of which completeness, internal consistency and plausibility are computed over the entire dataset while cross-source agreement is necessarily restricted to the race-years both sources hold.
R1-Rev2-Q8 — Event-level coverage
“Please clarify the completeness of the final dataset at the event level. Specifically, how many eligible IRONMAN® and IRONMAN® 70.3 race-years occurred during the study period, and what proportion of these are included. If there is no possibility to do so, this should be explicitly acknowledged as a limitation!”
Response. There is no possibility to do so, and we take the limitation the Reviewer explicitly offers. We want to be precise about why, because the reason is not merely that we did not look.
No authoritative public register exists of every IRONMAN® and IRONMAN® 70.3 race-year held since 2002, against which a coverage proportion could be computed. Moreover, our own event discovery was partly blocked by the official platform’s content delivery network, so the series list on which any internal denominator would rest is itself of unverified completeness. A proportion computed from it would look like a coverage rate while measuring only our own reach.
What we can report, and now do, is the pipeline in counts:
|
Event series identified |
128 |
|
Race editions enumerated across them |
1,235 |
|
Editions returning results |
1,172 (2,041,743 records) |
|
Editions returning no results |
63 |
|
Race-years recovered from the supplementary source |
382 (665,179 records) |
|
Distinct race-years in the merged dataset |
1,554 |
The 63 editions that returned no results are informative for anyone looking for a race-year they expect to find: 28 of them belong to the 2021 season, consistent with that year’s cancellations, and the remainder are scattered editions that were cancelled or whose results were never published.
We deliberately do not report a retrieval percentage. An earlier draft of this answer did, and it was wrong: the enumerated and retrieved sets were not nested, so the ratio described nothing. Correcting that led us to a defect in our own repository, reported in the final section of this letter.
Accordingly, Limitations now states:
“Population-scale” in this article therefore refers to the scale of athlete-race records assembled and not to a claim of event-level completeness.
R1-Rev2-Q9 — Were unavailable official race-years recovered via CoachCox?
“Please also clarify whether race-years unavailable from the official source were subsequently recovered through CoachCox portal.”
Response. Yes, and this is now quantified in Methods §Data merging. 382 race-years absent from the official source were recovered from the supplementary source, contributing 665,179 records — 24.6% of the dataset. The full three-way split is 559 race-years in both sources, 613 in the official source only, and 382 in the supplementary source only.
The supplement is therefore not a redundant second copy of the official data. Roughly a quarter of the dataset exists only because of it.
R1-Rev2-Q10 — Operational definitions of T1 and T2
“The operational definitions of T1 and T2 should be described in greater detail […] It may be influenced not only by athlete performance but also by transition-zone length, timing-mat positioning, congestion, race layout, and organizational procedures.”
Response. The Reviewer is describing a genuine limit on what these fields mean, and we have chosen to expose it rather than claim a uniformity we cannot support. A new Methods subsection, §Definition of the transition times, states:
Both platforms report T1 and T2 as intervals between timing mats: T1 between the swim exit and the start of the bike course, T2 between the end of the bike course and the start of the run. We neither defined these intervals nor positioned the mats; the fields are transcribed as published. Their content therefore depends on decisions taken by each event organizer — the size and layout of the transition area, where the mats sit relative to the bike racking, and how congested that area is when a given athlete passes through it. None of these are recorded in the results, and none are guaranteed to be constant across events or across years within the same event.
A transition time in this dataset is consequently the time an athlete took to traverse one particular transition area at one particular event, rather than a standardized measure of transition efficiency, and comparisons across events should be read with that in mind.
The Reviewer’s list appears again in Limitations, where the consequence is stated: a comparison of transition times between events carries an unmeasured component of course design.
R1-Rev2-Q11 — 2026 season incomplete
“The official and supplementary data were collected on 27–28 March 2026. Therefore, the reported description of the dataset as covering 2002–2026 requires clarification because the 2026 season is incomplete.”
Response. Accepted in full. We considered changing the coverage statement to 2002–2025 and decided against it: the 2026 records are valid and present, seven race-years from that season appear in the cross-source overlap, and discarding them would answer more than was asked. Instead, the partial status is marked everywhere the range appears:
- Abstract — “between 2002 and 2026, the final season partial”
- Methods — the collection dates are stated, followed by: “Because collection took place in March 2026, the 2026 season is represented only by races run before that date and is partial wherever it appears in this article.”
- Results and Figures 1, 4 and 5 — 2026 marked as a partial season in the captions and on the plots
- Limitations — listed as one of the three limitations of assembly
R1-Rev2-Q12 — How were the 6 validation races selected; are they representative?
“[…] only 6 races from the 2024 season were included. Please explain how these races were selected and why they are considered representative of a dataset covering more than two decades.”
Response. We are not able to give a defensible answer to this question as asked, and we prefer to say so rather than construct a sampling rationale after the fact. The six races were not selected by a stated rule.
The question is therefore answered by removal rather than by explanation. With the comparison rerun over all 559 overlapping race-years (Q1), there is no selection to justify and no representativeness to argue: the validation set is the entire overlap between the two sources, across 24 seasons and both race distances.
R1-Rev2-Q13 — Is CoachCox truly independent?
“If both sources ultimately originate from the same underlying official timing data, the observed agreement represents strong cross-source concordance, but not necessarily full independent external validation.”
Response. The Reviewer is right, and we concede it without qualification. The supplementary source aggregates results published by the same timing operation that supplies the official platform. Its independence is editorial, not metrological.
Methods §Data sources now states this directly:
Agreement between the two sources therefore establishes faithful transcription through two separate collection pipelines, and not independent measurement of the underlying times.
The same point is repeated in Limitations, so that a reader who arrives at the agreement figure from either direction meets the qualification with it.
We note that this concession and the expansion under Q1 belong together. We have given up the claim we could not support and strengthened the one we could.
R1-Rev2-Q14 — Test whether Unicode normalization improves the match rate
“The authors attribute the lower match rate to character-encoding differences. This explanation should be supported by an additional analysis whether the match rate improves (e.g. with Unicode normalization).”
Response. We ran the test the Reviewer asks for, and it showed that our stated explanation was incomplete. We report the result as it came out.
Matching was run three times over the whole overlap: case-insensitive exact; the same after Unicode NFKD normalisation, folding diacritics and punctuation; and an order-invariant rule that additionally sorts the name tokens. There are two distinct mechanisms, and encoding is the smaller one.
Name ordering, which we had not identified. Sixty-three race-years matched at essentially zero under exact matching despite record counts agreeing to within a few athletes. All are between 2017 and 2019 and all are full-distance. The supplementary source records names as “Lastname, Firstname” in those seasons while the official platform records “Firstname Lastname”. Diacritic folding cannot repair a reordering. Sorting the name tokens lifts those three seasons from near zero to 93.1%, 93.3% and 93.5%.
Character encoding, which is real but small. Unicode normalisation alone raises the median match rate by 0.5 percentage points. Where non-English names are frequent it contributes more — IRONMAN® Brazil 2024, the example our Table 2 caption cited, rises from 85.9% to 87.8% under normalisation and to 88.3% under the order-invariant rule — but that is roughly two points of a fourteen-point gap, not the whole of it.
The Table 2 caption generalised the encoding explanation to all cases. It has been rewritten to distinguish the two mechanisms, and the Results report both.
R1-Rev2-Q15 — How many authors performed the manual event-name checks?
“Please specify how many of the authors performed these checks and whether any independent verification was undertaken. If only one author conducted, then this should by reported.”
Response. One author performed the manual inspection of the normalised-name mapping, and it was not independently verified by a second. Methods §Data merging now states exactly that.
We note the checks were not the only safeguard on that step: the merge rule is deterministic and its counts are reproducible from the repository, so the 559 collisions and their resolution can be re-derived by a reader without relying on the manual pass.
R1-Rev2-Q16 — Equivalent missing-data assessment for T2
“The missing-data assessment currently focuses mainly on T1. Since separated T1 and T2 information is a central contribution of the dataset, an equivalent assessment for T2 should be provided, if possible.”
Response. There was no defensible reason to report one and not the other in a paper whose contribution is separated transition times. Every assessment previously reported for T1 has been repeated for T2.
Bias check. Among full-distance finishers, athletes without T1 data (n = 42,972) have a median overall finish of 11:46:20 against 12:27:08 for those with it (n = 1,051,808) — 41 minutes. The equivalent comparison for T2 gives a much smaller difference: 12:09:40 without (n = 25,706) against 12:25:59 with (n = 1,069,074) — 16 minutes. In both cases the records lacking the transition are concentrated in older races with faster, more competitive fields, and the effect is weaker for T2.
Coverage by year. T1 ranges from 62.6% (2006) to 89.2% (2024); T2 from 70.4% (2006) to 91.2% (2025). Both remain above 80% from 2008 onward.
And the asymmetry turned out to be informative rather than an artefact. Among finishers, T2 is better covered than T1 (98.3% against 95.8%). Among athletes who did not finish, the relation reverses sharply: T1 is present for 75.1% of DNF records, T2 for only 42.8%. This is what the race implies — an athlete who abandons on the bike leg passed through T1 and never reached T2 — and it means a missing T2 in a DNF record localizes the withdrawal rather than merely recording an absence. We would not have found this had the Reviewer not asked for the symmetric analysis.
Figure 3 now shows T2 alongside T1 in all four panels.
R1-Rev2-Q17 — Provenance of the plausibility cut-offs; thresholds for 70.3
“Please explain how these cut-offs were established and provide the corresponding thresholds for IRONMAN® 70.3.”
Response. Both parts are addressed, and the second uncovered a real defect.
Provenance. The bounds were set by the authors, from the observed distributions and from the physiological limits of the events. They were not adopted from a published source. The Methods now says so plainly; presenting author-chosen bounds as though they had external authority would have been worse than the vagueness it replaced.
The half-distance thresholds did not exist, and the check had never run on half-distance records. Applying the stated full-distance bounds to 70.3 finishers flags 85.1% on overall time, 54.2% on the bike split and 42.1% on the run — because the median IRONMAN® 70.3 finish is 5:53:16 against a stated lower bound of seven hours. Half-distance records are 50.5% of the dataset. The submitted claim that out-of-range values were below 1% for all fields therefore held only for the full distance.
We have derived distance-specific bounds for the half distance — swim 15–75 minutes, bike 1.5–5 hours, run 1–4 hours, overall 3.5–9 hours, with the same 0.5–30 minute transition range at either distance — and report the flag rates under them: below 0.7% for every field. The full-distance rates are below 1.2% for every field. Records outside the bounds for their distance are flagged, not removed.
R1-Rev2-Q18 — Table 1 coverage disagrees with §3.1.2
“The split-time coverage percentages in Table 1 do not appear to agree with those reported in 3.1.2 section. Please verify the calculations.”
Response. The Reviewer is correct, and the problem was larger than the four values the comparison revealed.
We recomputed every coverage figure from the merged file. The Results text is right; Table 1 was wrong in 20 of its 28 fields, not four. The rank fields were off by 19 to 21.6 percentage points — Table 1 gave rank_overall as 76.4% where the data gives 95.6%. bib was printed at 95.2% and is 100.0%. The three per-discipline distance fields were 6.4 to 7.4 points low.
This is not a denominator artefact; we checked against five candidate denominators and none reproduces the printed values. The cause is that Table 1 was typed once from an earlier data state and never regenerated after the final merge — the same cause as the figure defect reported under Q20.
We have addressed the mechanism, not only the instance. Table 1 and Table 2 are now produced by a script from audited result files, so a stale table cannot survive a rebuild. A second script verifies every numeric claim in the manuscript against the persisted results and exits non-zero on any mismatch; on its first run it failed 21 of 71 claims. It now passes 100 of 100.
We also note, for completeness, the correct split coverage: swim 85.9%, bike 85.9%, run 82.9%, overall 83.0%, T1 84.2%, T2 84.1%.
R1-Rev2-Q19 — “100%” within 5 s vs 57 records differing by >60 s
“The internal consistency section reported that 100% of records were within 5 seconds, while 57 records are subsequently reported to differ by more than 60 secs. Please report the exact percentage instead of rounding this to 100%.”
Response. Corrected. Over the 2,139,756 records where all five splits and the overall time are present and non-zero, 76.94% agree exactly, 98.30% fall within one second, and 99.9973% fall within five seconds. The 57 records are restated as 0.0027%.
The Reviewer’s point stands beyond this instance: a rounding that reads as a contradiction is a reporting error even when the underlying figure is right. We note that our numeric auditor cannot catch this class — it compares values against tolerances, and 99.9973 passes against a printed 100.0 — so rounding judgement stays a human check.
R1-Rev2-Q20 — Figure 1 content does not match its caption
“Figure 1 currently shown appears to contain split-time distributions, whereas the caption describes dataset composition by data source and race type. Please check.”
Response. We have checked, against the figure files in the submitted bundle. This particular observation does not hold: Figure 1 as submitted does show dataset composition, as its caption states, and the split-time distributions the Reviewer describes are Figure 2, also as captioned. We suspect the two were read in sequence.
Checking it, however, revealed an error in Figure 1 that we had not detected, and that neither reviewer reported. In panels (a) and (c), the two data sources were transposed. The bars labelled CoachCox carried the official platform’s records and vice versa — so the figure told a reader the supplementary source was three times the size of the primary one, contradicting the 75.4% / 24.6% split stated in the text on the same page. Panel (b) was correct.
We traced the cause: the plotting code assigned display labels to a grouped result by position, while the group order was set by the order of first appearance in the data rather than alphabetically. Panel (b) used a key-based rename and was therefore unaffected. Figure 1 has been regenerated, all figures now map labels by key with an assertion that fails on any unmapped category, and Figures 2 through 4 were re-audited for the same class of defect.
This defect and the Table 1 error under Q18 share one cause — derived artefacts generated from an earlier data state and never regenerated after the final merge. We report them together because they are one diagnosis, not two accidents.
We are grateful for the comment that led us to it.
R1-Rev2-Q21 — Figure 2 in minutes rather than hours
“Figure 2 presentation in minutes rather than hours would be easier to interpret for T1 and T2.”
Response. Done. The transition panels of Figure 2 are now in minutes; the swim, bike, run and overall panels remain in hours, where hours are the natural unit.
R1-Rev2-Q22 — Figure 3 shows only T1
“In Figure 3 only T1 coverage is presented, and T2 analysis should be added.”
Response. Done. All four panels of Figure 3 now show T2 alongside T1 — coverage by year, by data source, by race type, and by finish status. The fourth panel is where the DNF asymmetry described under Q16 becomes visible.
Panel (c) of Figure 1 likewise now shows both transitions by source.
R1-Rev2-Q23 — Add a demographic figure on male and female participation
“A demographic figure showing the temporal development of male and female participation would also add value for visual representation.”
Response. Added as Figure 5: annual record counts for male and female athletes, and the female share of records within each race distance, across the covered period with 2026 marked as partial.
The figure is referenced descriptively in Results §Demographics and carries no interpretive text there. We have deliberately not discussed what drives the trend: the dataset supports describing the composition, not explaining it, and the sex-difference literature we would need to interpret it against is cited among the research directions instead.
R1-Rev2-Q24 — Figure 4: mark 2026 as partial; label the pandemic period precisely
“[…] please clearly indicate that 2026 represents only a partial season. Also, pandemic period should also be labelled more precisely as the ‘COVID-19 pandemic period’.”
Response. Both done. Figure 4 marks 2026 as a partial season, and the shaded band is now labelled “COVID-19 pandemic period” in the panel and in the caption. The partial-season marking has been applied to Figures 1 and 5 as well, wherever the annual series runs to 2026.
R1-Rev2-Q25 — Greater attention to methodological limitations
“Well written, but greater attention should be given to methodological limitation. Examples: uncertainty regarding complete event-level coverage; whether the two sources are truly independent; differences in transition-zone design and timing procedures; incomplete coverage of the 2026 season; changes over time in races, courses, environmental conditions, and participant characteristics.”
Response. All five items are now in Limitations, each in the terms its own comment settled, and the section has been restructured rather than extended. It had become a chain of six ordinals and would have reached eleven; it is now organised into three passages — how the dataset was assembled, what the recorded values mean, and what the released file contains — so that each limitation is findable rather than buried in a list.
The five items map as follows:
|
Reviewer’s item |
Where it now sits |
Settled under |
|
Event-level coverage |
Assembly, first |
Q8 |
|
Source independence |
Assembly, second |
Q13 |
|
2026 partial |
Assembly, third |
Q11 |
|
Transition-zone design and timing procedures |
What the values mean |
Q10 |
|
Change over time in races, courses, conditions, participants |
What the values mean |
this comment and Rev1-Q9 |
R1-Rev2-Q26 — Integrate more relevant literature
“It is also recommended to integrate more relevant scientific literature into the discussion, but also into the introduction. At present, much of the section describes the dataset itself, while comparison with previous triathlon and endurance-performance research remains limited.”
Response. Addressed in the same pass as the Editor’s self-citation query, since the two point the same way — every external reference added to answer this comment lowers the ratio the Editor flagged.
Four external references were added: Laursen & Rhodes (2001) on why the determinants of endurance performance change beyond about four hours, Laursen et al. (2006) on thermoregulatory load measured during an IRONMAN® race, Millet & Lepers (2004) on progressive neuromuscular fatigue, and Jeukendrup (2011) on fuelling over multi-hour competition. They support the physiological-demands passage in the Introduction (Q4) and, in the Discussion, the grounds for stratifying analyses by race distance rather than pooling.
Three self-citations were removed as carrying no argument (see R1-Ed-1). The revised manuscript cites 17 references against 16, with the balance shifted from 8 self-citations to 5.
R1-Rev2-Q27 — Reconsider “validated” and “largest and most complete”
“The conclusion describes the resource as a validated dataset (…). These statements should be reconsidered after the methodological issues described above are addressed. Similarly, ‘largest and most complete (…)’ should either be objectively demonstrated or softened.”
Response. Reconsidered after addressing the issues, as the Reviewer instructs, and the two phrases resolve differently.
“Validated” is retained, on the basis set out under Q1: agreement between sources is now measured over 559 race-years and 4,732,776 discipline pairs rather than six races, alongside completeness, internal consistency and physiological plausibility computed over the whole dataset. We have also stated what the agreement does not establish (Q13), so the word is bounded rather than merely asserted.
“Most complete” is removed. “Largest openly available” is retained, with its comparator stated in the sentence.
The Conclusions now also report where the resource is weaker, in the same paragraph as its strengths: agreement is lower in the earliest seasons, transition times are absent for roughly one record in six, and the fields carried only by the official platform are unavailable for a quarter of the dataset.
Corrections we identified ourselves
Working through the comments surfaced errors that neither reviewer raised. Under Open Review we would rather report them than leave them to be found. Two are described above, in the responses to Q18 and Q20, because they share a cause with the points the reviewers made. Two more are recorded here.
The country count was not a count of countries. The submitted Results stated that athletes from 490 countries are represented. The country field is free text, and its distinct values include dropdown placeholders, US state abbreviations, bare initials, code fragments and misspellings; there are approximately 195 countries in the world, and any figure near 490 should have been implausible on its face. The claim is replaced by one computed from the ISO 3166-1 field: 251 distinct country codes, present for 74.2% of records. The leading-country shares were verified and are correct, and are retained from the free-text field because it covers 99.6% of records — the United States (937,613; 34.6%), the United Kingdom (183,093; 6.8%) and Australia (152,980; 5.7%). Their rank order was wrong in the submitted text and is corrected. The condition of the country field is now documented in Table 1 and in the reuse guidance, since it is exactly the kind of trap Rev1-Q7 asks us to warn reusers about.
Fifty-five records carried an upstream redaction we had not propagated. Fifty-five records hold a redaction marker in both the name and country fields. These come from the supplementary source, which honours suppression requests; the marker records a privacy decision taken by an athlete. Our deposit removes names for everyone, so the redacted name was invisible inside the general de-identification — but the redacted country would have survived into the deposited file, together with the event, age group, bib number and split times of people who had asked not to be identifiable. Those 55 records have been withdrawn from the deposit. The consequence is stated in the manuscript rather than smoothed over: the deposited file contains 2,706,867 records while the dataset described in this article comprises 2,706,922, and Table 1 marks which fields reach the deposit.
A new version of the Zenodo deposit (version 3.0.0, https://doi.org/10.5281/zenodo.22097794) has been published carrying the corrected file, the de-identified wording, and the full eight-author list. The Data Availability Statement now cites the concept DOI, which always resolves to the current version.
Withdrawing the records from the newest version was not by itself sufficient, and we mention this because it is the kind of half-measure that can pass for a fix. Zenodo versions are permanent — a published DOI must keep resolving and cannot be deleted — so the earlier version continued to distribute the file those 55 records were in, which is the version the submitted manuscript cited. Access to that file has now been closed. The earlier DOIs still resolve, and their record pages state that the version is superseded, why the file was withdrawn, and where the current data are.
Summary of changes
|
Comments addressed |
43 (1 editorial, 10 Reviewer 1, 27 Reviewer 2, 5 carried from the pre-check round) |
|
Comments rejected |
0 |
|
New analysis scripts |
11 |
|
Supporting scripts (figure and table generation, numeric audit) |
3 |
|
Figures regenerated |
4 |
|
Figures added |
1 |
|
Tables rebuilt from generated output |
2 |
|
References added / removed |
4 / 3 |
|
Self-citation rate |
50% → 29.4% |
|
Numeric claims verified against persisted results |
100 of 100 |
We thank the Editor and both Reviewers again. The manuscript is materially more accurate than the version they received, and in several places that is a direct consequence of questions they asked.
Reviewer 2 Report
Comments and Suggestions for AuthorsPlease find the attached document.
Comments for author File:
Comments.pdf
Author Response
Response to Reviewers — IRONMAN® Data Paper R1
Manuscript: “A Validated, Population-Scale Dataset of 2.7 Million IRONMAN® Triathlon Records with Separated Transition Times (2002–2026)” Manuscript ID: sci-4491984 Journal: Sci (MDPI) R1 decision received: 2026-08-19 R1 response date: [TO BE FILLED at resubmission]
Note on Open Review: this manuscript is under Open Review, so this response is published alongside the paper. It is written to be read by anyone, not only by the reviewers.
We thank the Editor and both Reviewers for reports that were specific enough to be acted on rather than merely acknowledged. Three of Reviewer 2’s comments asserted that we had made an error. We checked all three against the data. Two were correct, and checking the third — which does not reproduce — led us to a defect neither reviewer had seen. Working through the remaining comments surfaced two further errors that no reviewer raised. All four are reported below.
We have accepted every comment. There is no point on which we argue with a reviewer’s recommendation.
The substantive changes are:
- Cross-source validation expanded from 6 races to the entire overlap between the two sources — 559 race-years spanning 2003 to 2026, 340 full-distance and 219 half-distance, yielding 4,732,776 matched athlete-discipline pairs. This addresses Rev2-Q1 at its root: with the whole overlap compared, there is no sample to justify and no representativeness to argue.
- The independence of the second source is conceded (Rev2-Q13), and what the agreement statistic does and does not establish is now stated wherever it is reported.
- Eleven new analyses, among them T2 missingness mirroring T1, age-category stability across the covered years, half-distance plausibility bounds, merge-rule and event-coverage counts, the condition of the country field, and a three-rule athlete-matching comparison that revises an explanation we had given.
- Five figures regenerated and one added; Table 1 and Table 2 rebuilt from generated output rather than hand-maintained.
- A limitations section restructured around the five points Rev2-Q25
A note on how the numbers in this letter were produced. Every quantitative claim in the revised manuscript is now checked against a persisted result file by a script that runs after any regeneration and exits non-zero on a mismatch; it currently verifies 100 claims. We built it during this revision because two of the errors we found were caused by figures that had been correct when first computed and were never recomputed afterwards.
Editor’s Letter
R1-Ed-1 — Self-citation rate
“During the technical check of your manuscript, we noticed that a high proportion of the cited references belong to you or your co-authors Refs. 2,3,4,5,7,12,16, which is a self-citation rate of about 43.75%. Could you please check whether the inclusion of each of these references is appropriate?”
Response. We checked, and the check produced an answer we did not expect: the rate was higher than the office reported.
Reference 15 (Loosli et al., Frontiers in Sports and Active Living, 2025) is also a self-citation — Nikolaidis, Andrade, Rosemann and Knechtle are among its ten authors — and was not flagged. The reason it was not flagged is our own error: the bibliography entry truncated the author list, so the reference rendered under the first author alone and the four co-authors were invisible to anyone counting. The submitted manuscript’s true self-citation rate was therefore 8 of 16 (50%), not 43.75%. We have corrected the entry, which also corrects the first author’s forename.
We then mapped each flagged reference to the claim it supports and classified it as load-bearing or not. Reference numbers below are those of the submitted version, which is what the office counted; restructuring the Introduction has since changed the order of first appearance and therefore the numbering.
Removed — three references that carried no argument:
|
Ref |
Reason |
|
4 — Knechtle et al. 2025 (environment) |
Supported the same Introduction sentence as reference 3, on the same point. One citation suffices. |
|
12 — Knechtle et al. 2024 (sex differences) |
A single mention inside an enumeration of possible research directions. The enumeration lists topics; it does not argue from sources. |
|
16 — Sousa et al. 2021 |
The same, in the same enumeration. |
Retained — each with a claim that fails without it:
|
Ref |
The claim it carries |
|
2 — Nikolaidis et al. 2023 |
The comparator for the dataset’s size. The statement that this resource is 3.3 times larger than the largest previously published IRONMAN® analysis requires the analysis it is measured against (823,459 records). |
|
3 — Knechtle et al. 2025 (origin) |
The scale of prior full-distance work (677,320 records), which establishes what “population-scale” meant before this dataset. |
|
5 — Knechtle et al. 2019 |
The only prior population-scale measurement of transition time we are aware of — that the fastest finishers spend 0.9% of race time in transitions against 2.2% for the slowest. It is the quantitative basis for treating transitions as worth separating. |
|
7 — Rüst et al. 2014 |
The study this dataset exists to answer. It is the one dedicated analysis of transition times at IRONMAN® distance we are aware of, it was restricted to the annual top ten World Championship finishers, and it could not separate T1 from T2. Removing it would remove the paper’s motivation. |
|
15 — Loosli et al. 2025 |
The review behind the statement that the research directions we list are recognised gaps rather than our own suggestions. |
We have simultaneously added four external references, in answering Rev2-Q4 and Rev2-Q26: Laursen & Rhodes (2001), Laursen et al. (2006), Millet & Lepers (2004) and Jeukendrup (2011).
The revised manuscript cites 17 references, of which 5 are self-citations — 29.4%, against a true submitted rate of 50%.
Reviewer 1
We thank Reviewer 1 for a reading that identified where the methods section was thin. Most of the ten points asked us to state something we knew but had not written down, and in two cases (Q6, Q8) writing it down required computing it, which we had not done.
R1-Rev1-Q1 — Abstract opening sentence carries two messages
“The first sentence is quite long and conveys two different messages […] the first sentence should be split into two shorter sentences so that the gaps in the literature are addressed first, followed by a clear statement of the study’s objective.”
Response. Done, in the order the Reviewer specifies. The Abstract now opens with the field’s use of long-distance triathlon as a model, states the constraint that existing datasets impose as a separate sentence, and follows with the aim. Splitting the opening sentence and adding an explicit aim did not lengthen the Abstract: it is 205 words against 207 in the submitted version.
R1-Rev1-Q2 — State the purpose in separate, clear sentences
“The purpose of the study must be stated in separate and clear sentences.”
Response. An explicit aim sentence has been added as the third sentence of the Abstract: “Our aim was to assemble, validate, and describe an openly available dataset that removes both constraints.” The wording was chosen after settling the position on “validated” that Rev2-Q1 raises, so that the two are consistent.
R1-Rev1-Q3 — Validation results as three consecutive percentages
“Instead, the summary could begin with a general conclusion […] The most important validation findings could then be listed.”
Response. Adopted. The Abstract now leads with the conclusion — “The dataset showed close correspondence between sources and high internal consistency” — and the figures follow it. The figures themselves changed, because the validation was rerun over a far larger comparison (Rev2-Q1) and because one of them was a rounding that read as a contradiction (Rev2-Q19); the Abstract carries the corrected values.
R1-Rev1-Q4 — Introduction final paragraph overpromises applications
“Some examples—such as examining post-COVID-19 return-to-competition patterns or the impact of shoe technology on running performance—give the impression that the dataset contains variables that would directly support these analyses.”
Response. The Reviewer is right, and we have taken the sharper of the two available routes. Rather than hedging each example, the research directions are now explicitly divided into those the schema supports and those it does not:
Others would require linkage to information the dataset does not hold: return-to-competition patterns after the COVID-19 disruption can be described in aggregate but not attributed, since no field records why an athlete was absent, and any study of advanced footwear technology would need external data on what athletes wore, which no results platform publishes. We list the second group to mark the boundary rather than to claim it.
The paragraph also moved. It had been duplicated between the Introduction and the Discussion; it now appears once, in Discussion §Research directions, which is what Rev2-Q7 asks for from the other direction.
R1-Rev1-Q5 — No clear research objective at the end of the Introduction
“[…] specifying a clear research objective and/or research question at the end of the introduction would strengthen the focus of the study.”
Response. An aim sentence was present in the submitted Introduction, but both Reviewers missed it — which told us the problem was placement rather than absence. It sat before a block of methodological description that Rev2-Q7 asks to be moved out. Moving that block leaves the aim as the closing statement of the Introduction, which is where both Reviewers looked for it.
R1-Rev1-Q6 — How duplicates across sources are identified and removed
“[…] if records for the same athlete or the same race appear in both sources, a more detailed explanation is needed regarding how duplicate records are identified and the method used to remove them.”
Response. The rule is now written out in Methods §Data merging, with the counts it produces:
- Deduplication is at the race-year level, not the athlete level. Overlap is detected by normalising event names — stripping year, brand prefixes and championship designations — and matching on normalised name, year and race type.
- When a race-year is present in both sources, the official record is retained and the supplementary record discarded, without exception.
- Applied across the two sources, 559 race-years were present in both and resolved in favour of the official source in every case, against 613 present only in the official source and 382 present only in the supplementary source — 1,554 distinct race-years in the merged dataset.
We also report what the normalisation contributes, because it is the step the rule depends on: matching raw event names identifies none of the 559 collisions, the generic rules identify 537, and the championship-prefix rules account for the remaining 22.
R1-Rev1-Q7 — Rationale for combining IRONMAN® and IRONMAN® 70.3
“The rationale for combining IRONMAN® and IRONMAN® 70.3 races into the same database should be explained more clearly […] it would be helpful to provide recommendations regarding distance-based analyses.”
Response. Both parts are now addressed, in the two places they belong.
The rationale is in Methods §Data consolidation: a shared schema is precisely what makes direct comparison across distances possible, and distance-specific datasets preclude it by construction. The race_type field preserves the distinction inside the file, so pooling remains a decision the analyst takes rather than one the data imposes.
The recommendation is in Discussion §Practical guidance for reuse, and we have given it the quantitative grounds the Reviewer’s own argument implies. The median finish among finishers is 5:53:16 at the half distance against 12:25:37 at the full distance, and the determinants of endurance performance are not the same across a range that wide (Laursen & Rhodes, 2001). Analyses should stratify by race_type or justify pooling explicitly.
R1-Rev1-Q8 — Did age categories change over the covered years?
“[…] it should be clarified whether any changes were made to the age categories of the competitors during those years.”
Response. We did not know, so we checked rather than asserting stability. Distinct age-group values were tabulated by race year across the whole dataset.
No boundary moves at any point. The five-year bands from 18-24 to 75-79 are present in every year from 2002 to 2026. What varies is only which of the oldest bands is populated: 80-84 is absent from 2006 to 2008, and 85-89 appears from 2015 to 2018 and again from 2022 onward. This reflects whether any athlete in those bands finished a given season, not a change in the classification. The finding is reported in Methods §Data consolidation.
R1-Rev1-Q9 — Environmental variables absent from the dataset
“It should be clearly stated that the dataset does not include environmental variables—such as course profile, elevation, air and water temperature, wind conditions, and drafting rules […]”
Response. Stated in Limitations, in the Reviewer’s own terms:
The dataset holds no course profile, elevation, air or water temperature, wind, or drafting regulation, so analyses of performance across events cannot control for the conditions under which those performances were produced.
We have added the temporal counterpart, which Reviewer 2 raises as the fifth item of Rev2-Q25: courses, qualification rules and participant populations changed over the 24 years covered, which makes any longitudinal comparison a comparison between eras as much as between athletes.
R1-Rev1-Q10 — Justify or soften “largest and most comprehensive”
“[…] it should be made clearer on what criteria strong statements such as ‘the largest and most comprehensive dataset in the field’ are based, or such statements should be phrased more cautiously.”
Response. We have separated the two claims, because they are not equally supportable.
“Largest” is retained with its comparator stated. At 2,706,922 records the dataset is more than three times the size of the largest previously published IRONMAN® analysis (823,459 records, Nikolaidis et al. 2023), and it spans both race distances rather than one. That is a checkable statement and it now appears with the number it is measured against, in the Discussion and in the Conclusions.
“Most comprehensive” and “most complete” are removed. They are the weaker claims, they are the ones Rev2-Q8 attacks from the coverage side, and we cannot demonstrate them — as the answer to Rev2-Q8 explains, no register exists against which event-level completeness could be computed. The Conclusions now name the resource’s weaknesses alongside its strengths.
Reviewer 2
We thank Reviewer 2 for a report that identified the manuscript’s central weakness correctly. The opening observation — that the reporting issues “limit the interpretation of the dataset as fully ‘validated’ and ‘population-scale’” — is one argument arriving from several directions, and we have answered it as one argument rather than hedging each mention separately.
Our position, stated once and referenced below. “Validated” is retained, and earned rather than defended: the cross-source comparison has been rerun over the entire overlap between the two sources instead of a six-race sample. “Population-scale” is retained with its meaning narrowed to what we can support — the scale of records assembled — and the event-level completeness the term might imply is declared unquantifiable, which is the route the Reviewer explicitly offered.
R1-Rev2-Q1 — “Validated” too strong for 6 races from one season
“[…] the term ‘validated’ appears relatively strong considering that cross-source validation was performed on 6 races from one season. The authors should provide stronger justification for this terminology or consider more cautious wording […]”
Response. We chose the first of the two routes the Reviewer offers, and provided the justification by enlarging the evidence rather than by arguing about the existing evidence.
The submitted validation used 6 races, all from the 2024 season. A census of the two raw sources showed that 559 race-years are present in both, spanning 2003 to 2026, 340 full-distance and 219 half-distance, with roughly one million records available for pairing on each side. The comparison has been rerun over all of them.
Results, now in Results §Technical validation and Table 2:
- Record counts from the two sources differ by two or fewer athletes for 484 of the 559 race-years (86.6%).
- Athlete matching yields a median rate of 3% per race-year under case-insensitive exact matching and 96.4% under an order-invariant rule.
- Among matched athletes, 4,732,776 athlete-discipline pairs were compared across all six disciplines. 76% agree exactly, to the second, ranging from 98.50% for the run split to 99.10% for the swim.
- Agreement is not uniform over time, and we report this rather than pooling it away: the median race-year agrees exactly for 100% of pairs and the fifth percentile for 99.21%, but 2017 sits at 86.3% and 2014 at 95.1%, against 99.9% or above from 2020 onward.
One caution, which we state in the Methods and repeat here. The 98.76% figure is not a like-for-like revision of the 99.7–99.9% reported in the submitted version. It covers 559 race-years instead of 6, and it is computed on an order-invariant match, which is a more permissive rule admitting pairs the earlier exact match never saw — including the 2017–2019 seasons, which matched at essentially zero before. A reader comparing 99.8% to 98.8% would conclude the data had got worse. The measurement got broader.
R1-Rev2-Q2 — “Five continents” may be four
“This statement that the 6 validation races represented five continents should also be checked. Based on the events listed in the methods, only 4 continents appear to be represented.”
Response. The Reviewer is correct. The six races were Florida, Frankfurt, Brazil, South Africa, 70.3 Oceanside and 70.3 Louisville — four distinct continents, with three of the six in North America. The claim appeared in four places: the Abstract, Methods, Results and the Table 2 caption.
We have not corrected the number, because the sentences containing it no longer exist. The expanded comparison under Q1 replaced them, and the validation set is now described by what it is — every race-year both sources hold, across 24 seasons and both distances — rather than by a continent count.
R1-Rev2-Q3 — Add IRONMAN to the keywords
“Please consider adding IRONMAN to the keywords”
Response. Added. The keyword list now reads: triathlon; IRONMAN®; open dataset; transition times; data validation; endurance performance; race results; reproducibility. The registered trademark symbol is carried as it is on every other mention in the manuscript.
R1-Rev2-Q4 — “Extreme physiological demands” not described
“[…] the manuscript refers to the ‘extreme physiological demands’ of IRONMAN® triathlon, but these demands are not further described. […] This does not need to become an extensive physiological review; one or two well-supported sentences would be sufficient.”
Response. Two sentences added to the Introduction, and no more, per the Reviewer’s own constraint:
Those demands are of a kind that the usual determinants of endurance performance do not fully capture: beyond roughly four hours of competition, maximal oxygen uptake and the anaerobic threshold cease to predict performance well, and fuel and fluid provision, substrate availability, and electrolyte balance become limiting in their own right [Laursen & Rhodes 2001; Jeukendrup 2011]. Competition over eight to seventeen hours also imposes a sustained thermoregulatory load [Laursen et al. 2006] and a progressive loss of neuromuscular function across the three disciplines [Millet & Lepers 2004], so that the athlete who begins the run is not physiologically the athlete who began the swim.
All four references are external to the author list, which also serves the Editor’s query.
R1-Rev2-Q5 — Rationale for separating T1 from T2
“[…] the importance of separating T1 and T2 should be explained more clearly. […] T1 occurs after swimming and before cycling, whereas T2 follows prolonged cycling before running.”
Response. The asymmetry the Reviewer describes is now stated in the Introduction as the reason the separation matters:
T1 follows the swim and precedes the cycle, so it is dominated by the change of equipment and by the shift from horizontal to upright posture. T2 follows several hours of cycling and precedes the run, placing it at the point where the run-off-bike penalty originates; it therefore reflects accumulated fatigue as much as logistical efficiency. Treating them as a single combined interval discards that distinction.
The point is not only conceptual, and the dataset now demonstrates it. Among athletes who did not finish, T1 is present for 75.1% of records but T2 for only 42.8% — because an athlete who abandons on the bike leg has passed through T1 and never reaches T2. A combined interval cannot carry that information.
R1-Rev2-Q6 — “Only” and “first” used cautiously
“[…] strong novelty statements such as ‘only’ and ‘first’ should be used cautiously unless supported by a comprehensive search.”
Response. We considered documenting a search to support the claims, and did not, because no systematic search was performed and describing one after the fact would not be honest. The claims are softened instead, to exactly what we can stand behind:
- “The one dedicated analysis of transition times at IRONMAN® distance that we are aware of […]”
- “We are not aware of a peer-reviewed, openly available, and formally validated dataset for IRONMAN® triathlon […]”
We have kept the corresponding claim in the Conclusions in the same register, and the one claim that is checkable — that this is the largest openly available resource of its kind — is stated with its comparator rather than hedged.
R1-Rev2-Q7 — Methodological content in the Introduction
“[…] the final part contains substantial methodological information. Most of this would fit better in the materials and methods section. The introduction should preferably conclude with a clear knowledge gap and concise study aims.”
Response. Done, and the block turned out to need deleting rather than moving. Its content — the dataset’s composition and a list of potential applications — duplicated Results §Dataset composition and Discussion §Research directions almost sentence for sentence. Relocating it to the Methods would have created a third copy. The Introduction now closes on the knowledge gap and the aim.
This single edit also answers Rev1-Q4 and Rev1-Q5.
Separately, the Methods gained an explicit statement of the quality-assessment structure, which the submitted version left implicit: four dimensions, of which completeness, internal consistency and plausibility are computed over the entire dataset while cross-source agreement is necessarily restricted to the race-years both sources hold.
R1-Rev2-Q8 — Event-level coverage
“Please clarify the completeness of the final dataset at the event level. Specifically, how many eligible IRONMAN® and IRONMAN® 70.3 race-years occurred during the study period, and what proportion of these are included. If there is no possibility to do so, this should be explicitly acknowledged as a limitation!”
Response. There is no possibility to do so, and we take the limitation the Reviewer explicitly offers. We want to be precise about why, because the reason is not merely that we did not look.
No authoritative public register exists of every IRONMAN® and IRONMAN® 70.3 race-year held since 2002, against which a coverage proportion could be computed. Moreover, our own event discovery was partly blocked by the official platform’s content delivery network, so the series list on which any internal denominator would rest is itself of unverified completeness. A proportion computed from it would look like a coverage rate while measuring only our own reach.
What we can report, and now do, is the pipeline in counts:
|
Event series identified |
128 |
|
Race editions enumerated across them |
1,235 |
|
Editions returning results |
1,172 (2,041,743 records) |
|
Editions returning no results |
63 |
|
Race-years recovered from the supplementary source |
382 (665,179 records) |
|
Distinct race-years in the merged dataset |
1,554 |
The 63 editions that returned no results are informative for anyone looking for a race-year they expect to find: 28 of them belong to the 2021 season, consistent with that year’s cancellations, and the remainder are scattered editions that were cancelled or whose results were never published.
We deliberately do not report a retrieval percentage. An earlier draft of this answer did, and it was wrong: the enumerated and retrieved sets were not nested, so the ratio described nothing. Correcting that led us to a defect in our own repository, reported in the final section of this letter.
Accordingly, Limitations now states:
“Population-scale” in this article therefore refers to the scale of athlete-race records assembled and not to a claim of event-level completeness.
R1-Rev2-Q9 — Were unavailable official race-years recovered via CoachCox?
“Please also clarify whether race-years unavailable from the official source were subsequently recovered through CoachCox portal.”
Response. Yes, and this is now quantified in Methods §Data merging. 382 race-years absent from the official source were recovered from the supplementary source, contributing 665,179 records — 24.6% of the dataset. The full three-way split is 559 race-years in both sources, 613 in the official source only, and 382 in the supplementary source only.
The supplement is therefore not a redundant second copy of the official data. Roughly a quarter of the dataset exists only because of it.
R1-Rev2-Q10 — Operational definitions of T1 and T2
“The operational definitions of T1 and T2 should be described in greater detail […] It may be influenced not only by athlete performance but also by transition-zone length, timing-mat positioning, congestion, race layout, and organizational procedures.”
Response. The Reviewer is describing a genuine limit on what these fields mean, and we have chosen to expose it rather than claim a uniformity we cannot support. A new Methods subsection, §Definition of the transition times, states:
Both platforms report T1 and T2 as intervals between timing mats: T1 between the swim exit and the start of the bike course, T2 between the end of the bike course and the start of the run. We neither defined these intervals nor positioned the mats; the fields are transcribed as published. Their content therefore depends on decisions taken by each event organizer — the size and layout of the transition area, where the mats sit relative to the bike racking, and how congested that area is when a given athlete passes through it. None of these are recorded in the results, and none are guaranteed to be constant across events or across years within the same event.
A transition time in this dataset is consequently the time an athlete took to traverse one particular transition area at one particular event, rather than a standardized measure of transition efficiency, and comparisons across events should be read with that in mind.
The Reviewer’s list appears again in Limitations, where the consequence is stated: a comparison of transition times between events carries an unmeasured component of course design.
R1-Rev2-Q11 — 2026 season incomplete
“The official and supplementary data were collected on 27–28 March 2026. Therefore, the reported description of the dataset as covering 2002–2026 requires clarification because the 2026 season is incomplete.”
Response. Accepted in full. We considered changing the coverage statement to 2002–2025 and decided against it: the 2026 records are valid and present, seven race-years from that season appear in the cross-source overlap, and discarding them would answer more than was asked. Instead, the partial status is marked everywhere the range appears:
- Abstract — “between 2002 and 2026, the final season partial”
- Methods — the collection dates are stated, followed by: “Because collection took place in March 2026, the 2026 season is represented only by races run before that date and is partial wherever it appears in this article.”
- Results and Figures 1, 4 and 5 — 2026 marked as a partial season in the captions and on the plots
- Limitations — listed as one of the three limitations of assembly
R1-Rev2-Q12 — How were the 6 validation races selected; are they representative?
“[…] only 6 races from the 2024 season were included. Please explain how these races were selected and why they are considered representative of a dataset covering more than two decades.”
Response. We are not able to give a defensible answer to this question as asked, and we prefer to say so rather than construct a sampling rationale after the fact. The six races were not selected by a stated rule.
The question is therefore answered by removal rather than by explanation. With the comparison rerun over all 559 overlapping race-years (Q1), there is no selection to justify and no representativeness to argue: the validation set is the entire overlap between the two sources, across 24 seasons and both race distances.
R1-Rev2-Q13 — Is CoachCox truly independent?
“If both sources ultimately originate from the same underlying official timing data, the observed agreement represents strong cross-source concordance, but not necessarily full independent external validation.”
Response. The Reviewer is right, and we concede it without qualification. The supplementary source aggregates results published by the same timing operation that supplies the official platform. Its independence is editorial, not metrological.
Methods §Data sources now states this directly:
Agreement between the two sources therefore establishes faithful transcription through two separate collection pipelines, and not independent measurement of the underlying times.
The same point is repeated in Limitations, so that a reader who arrives at the agreement figure from either direction meets the qualification with it.
We note that this concession and the expansion under Q1 belong together. We have given up the claim we could not support and strengthened the one we could.
R1-Rev2-Q14 — Test whether Unicode normalization improves the match rate
“The authors attribute the lower match rate to character-encoding differences. This explanation should be supported by an additional analysis whether the match rate improves (e.g. with Unicode normalization).”
Response. We ran the test the Reviewer asks for, and it showed that our stated explanation was incomplete. We report the result as it came out.
Matching was run three times over the whole overlap: case-insensitive exact; the same after Unicode NFKD normalisation, folding diacritics and punctuation; and an order-invariant rule that additionally sorts the name tokens. There are two distinct mechanisms, and encoding is the smaller one.
Name ordering, which we had not identified. Sixty-three race-years matched at essentially zero under exact matching despite record counts agreeing to within a few athletes. All are between 2017 and 2019 and all are full-distance. The supplementary source records names as “Lastname, Firstname” in those seasons while the official platform records “Firstname Lastname”. Diacritic folding cannot repair a reordering. Sorting the name tokens lifts those three seasons from near zero to 93.1%, 93.3% and 93.5%.
Character encoding, which is real but small. Unicode normalisation alone raises the median match rate by 0.5 percentage points. Where non-English names are frequent it contributes more — IRONMAN® Brazil 2024, the example our Table 2 caption cited, rises from 85.9% to 87.8% under normalisation and to 88.3% under the order-invariant rule — but that is roughly two points of a fourteen-point gap, not the whole of it.
The Table 2 caption generalised the encoding explanation to all cases. It has been rewritten to distinguish the two mechanisms, and the Results report both.
R1-Rev2-Q15 — How many authors performed the manual event-name checks?
“Please specify how many of the authors performed these checks and whether any independent verification was undertaken. If only one author conducted, then this should by reported.”
Response. One author performed the manual inspection of the normalised-name mapping, and it was not independently verified by a second. Methods §Data merging now states exactly that.
We note the checks were not the only safeguard on that step: the merge rule is deterministic and its counts are reproducible from the repository, so the 559 collisions and their resolution can be re-derived by a reader without relying on the manual pass.
R1-Rev2-Q16 — Equivalent missing-data assessment for T2
“The missing-data assessment currently focuses mainly on T1. Since separated T1 and T2 information is a central contribution of the dataset, an equivalent assessment for T2 should be provided, if possible.”
Response. There was no defensible reason to report one and not the other in a paper whose contribution is separated transition times. Every assessment previously reported for T1 has been repeated for T2.
Bias check. Among full-distance finishers, athletes without T1 data (n = 42,972) have a median overall finish of 11:46:20 against 12:27:08 for those with it (n = 1,051,808) — 41 minutes. The equivalent comparison for T2 gives a much smaller difference: 12:09:40 without (n = 25,706) against 12:25:59 with (n = 1,069,074) — 16 minutes. In both cases the records lacking the transition are concentrated in older races with faster, more competitive fields, and the effect is weaker for T2.
Coverage by year. T1 ranges from 62.6% (2006) to 89.2% (2024); T2 from 70.4% (2006) to 91.2% (2025). Both remain above 80% from 2008 onward.
And the asymmetry turned out to be informative rather than an artefact. Among finishers, T2 is better covered than T1 (98.3% against 95.8%). Among athletes who did not finish, the relation reverses sharply: T1 is present for 75.1% of DNF records, T2 for only 42.8%. This is what the race implies — an athlete who abandons on the bike leg passed through T1 and never reached T2 — and it means a missing T2 in a DNF record localizes the withdrawal rather than merely recording an absence. We would not have found this had the Reviewer not asked for the symmetric analysis.
Figure 3 now shows T2 alongside T1 in all four panels.
R1-Rev2-Q17 — Provenance of the plausibility cut-offs; thresholds for 70.3
“Please explain how these cut-offs were established and provide the corresponding thresholds for IRONMAN® 70.3.”
Response. Both parts are addressed, and the second uncovered a real defect.
Provenance. The bounds were set by the authors, from the observed distributions and from the physiological limits of the events. They were not adopted from a published source. The Methods now says so plainly; presenting author-chosen bounds as though they had external authority would have been worse than the vagueness it replaced.
The half-distance thresholds did not exist, and the check had never run on half-distance records. Applying the stated full-distance bounds to 70.3 finishers flags 85.1% on overall time, 54.2% on the bike split and 42.1% on the run — because the median IRONMAN® 70.3 finish is 5:53:16 against a stated lower bound of seven hours. Half-distance records are 50.5% of the dataset. The submitted claim that out-of-range values were below 1% for all fields therefore held only for the full distance.
We have derived distance-specific bounds for the half distance — swim 15–75 minutes, bike 1.5–5 hours, run 1–4 hours, overall 3.5–9 hours, with the same 0.5–30 minute transition range at either distance — and report the flag rates under them: below 0.7% for every field. The full-distance rates are below 1.2% for every field. Records outside the bounds for their distance are flagged, not removed.
R1-Rev2-Q18 — Table 1 coverage disagrees with §3.1.2
“The split-time coverage percentages in Table 1 do not appear to agree with those reported in 3.1.2 section. Please verify the calculations.”
Response. The Reviewer is correct, and the problem was larger than the four values the comparison revealed.
We recomputed every coverage figure from the merged file. The Results text is right; Table 1 was wrong in 20 of its 28 fields, not four. The rank fields were off by 19 to 21.6 percentage points — Table 1 gave rank_overall as 76.4% where the data gives 95.6%. bib was printed at 95.2% and is 100.0%. The three per-discipline distance fields were 6.4 to 7.4 points low.
This is not a denominator artefact; we checked against five candidate denominators and none reproduces the printed values. The cause is that Table 1 was typed once from an earlier data state and never regenerated after the final merge — the same cause as the figure defect reported under Q20.
We have addressed the mechanism, not only the instance. Table 1 and Table 2 are now produced by a script from audited result files, so a stale table cannot survive a rebuild. A second script verifies every numeric claim in the manuscript against the persisted results and exits non-zero on any mismatch; on its first run it failed 21 of 71 claims. It now passes 100 of 100.
We also note, for completeness, the correct split coverage: swim 85.9%, bike 85.9%, run 82.9%, overall 83.0%, T1 84.2%, T2 84.1%.
R1-Rev2-Q19 — “100%” within 5 s vs 57 records differing by >60 s
“The internal consistency section reported that 100% of records were within 5 seconds, while 57 records are subsequently reported to differ by more than 60 secs. Please report the exact percentage instead of rounding this to 100%.”
Response. Corrected. Over the 2,139,756 records where all five splits and the overall time are present and non-zero, 76.94% agree exactly, 98.30% fall within one second, and 99.9973% fall within five seconds. The 57 records are restated as 0.0027%.
The Reviewer’s point stands beyond this instance: a rounding that reads as a contradiction is a reporting error even when the underlying figure is right. We note that our numeric auditor cannot catch this class — it compares values against tolerances, and 99.9973 passes against a printed 100.0 — so rounding judgement stays a human check.
R1-Rev2-Q20 — Figure 1 content does not match its caption
“Figure 1 currently shown appears to contain split-time distributions, whereas the caption describes dataset composition by data source and race type. Please check.”
Response. We have checked, against the figure files in the submitted bundle. This particular observation does not hold: Figure 1 as submitted does show dataset composition, as its caption states, and the split-time distributions the Reviewer describes are Figure 2, also as captioned. We suspect the two were read in sequence.
Checking it, however, revealed an error in Figure 1 that we had not detected, and that neither reviewer reported. In panels (a) and (c), the two data sources were transposed. The bars labelled CoachCox carried the official platform’s records and vice versa — so the figure told a reader the supplementary source was three times the size of the primary one, contradicting the 75.4% / 24.6% split stated in the text on the same page. Panel (b) was correct.
We traced the cause: the plotting code assigned display labels to a grouped result by position, while the group order was set by the order of first appearance in the data rather than alphabetically. Panel (b) used a key-based rename and was therefore unaffected. Figure 1 has been regenerated, all figures now map labels by key with an assertion that fails on any unmapped category, and Figures 2 through 4 were re-audited for the same class of defect.
This defect and the Table 1 error under Q18 share one cause — derived artefacts generated from an earlier data state and never regenerated after the final merge. We report them together because they are one diagnosis, not two accidents.
We are grateful for the comment that led us to it.
R1-Rev2-Q21 — Figure 2 in minutes rather than hours
“Figure 2 presentation in minutes rather than hours would be easier to interpret for T1 and T2.”
Response. Done. The transition panels of Figure 2 are now in minutes; the swim, bike, run and overall panels remain in hours, where hours are the natural unit.
R1-Rev2-Q22 — Figure 3 shows only T1
“In Figure 3 only T1 coverage is presented, and T2 analysis should be added.”
Response. Done. All four panels of Figure 3 now show T2 alongside T1 — coverage by year, by data source, by race type, and by finish status. The fourth panel is where the DNF asymmetry described under Q16 becomes visible.
Panel (c) of Figure 1 likewise now shows both transitions by source.
R1-Rev2-Q23 — Add a demographic figure on male and female participation
“A demographic figure showing the temporal development of male and female participation would also add value for visual representation.”
Response. Added as Figure 5: annual record counts for male and female athletes, and the female share of records within each race distance, across the covered period with 2026 marked as partial.
The figure is referenced descriptively in Results §Demographics and carries no interpretive text there. We have deliberately not discussed what drives the trend: the dataset supports describing the composition, not explaining it, and the sex-difference literature we would need to interpret it against is cited among the research directions instead.
R1-Rev2-Q24 — Figure 4: mark 2026 as partial; label the pandemic period precisely
“[…] please clearly indicate that 2026 represents only a partial season. Also, pandemic period should also be labelled more precisely as the ‘COVID-19 pandemic period’.”
Response. Both done. Figure 4 marks 2026 as a partial season, and the shaded band is now labelled “COVID-19 pandemic period” in the panel and in the caption. The partial-season marking has been applied to Figures 1 and 5 as well, wherever the annual series runs to 2026.
R1-Rev2-Q25 — Greater attention to methodological limitations
“Well written, but greater attention should be given to methodological limitation. Examples: uncertainty regarding complete event-level coverage; whether the two sources are truly independent; differences in transition-zone design and timing procedures; incomplete coverage of the 2026 season; changes over time in races, courses, environmental conditions, and participant characteristics.”
Response. All five items are now in Limitations, each in the terms its own comment settled, and the section has been restructured rather than extended. It had become a chain of six ordinals and would have reached eleven; it is now organised into three passages — how the dataset was assembled, what the recorded values mean, and what the released file contains — so that each limitation is findable rather than buried in a list.
The five items map as follows:
|
Reviewer’s item |
Where it now sits |
Settled under |
|
Event-level coverage |
Assembly, first |
Q8 |
|
Source independence |
Assembly, second |
Q13 |
|
2026 partial |
Assembly, third |
Q11 |
|
Transition-zone design and timing procedures |
What the values mean |
Q10 |
|
Change over time in races, courses, conditions, participants |
What the values mean |
this comment and Rev1-Q9 |
R1-Rev2-Q26 — Integrate more relevant literature
“It is also recommended to integrate more relevant scientific literature into the discussion, but also into the introduction. At present, much of the section describes the dataset itself, while comparison with previous triathlon and endurance-performance research remains limited.”
Response. Addressed in the same pass as the Editor’s self-citation query, since the two point the same way — every external reference added to answer this comment lowers the ratio the Editor flagged.
Four external references were added: Laursen & Rhodes (2001) on why the determinants of endurance performance change beyond about four hours, Laursen et al. (2006) on thermoregulatory load measured during an IRONMAN® race, Millet & Lepers (2004) on progressive neuromuscular fatigue, and Jeukendrup (2011) on fuelling over multi-hour competition. They support the physiological-demands passage in the Introduction (Q4) and, in the Discussion, the grounds for stratifying analyses by race distance rather than pooling.
Three self-citations were removed as carrying no argument (see R1-Ed-1). The revised manuscript cites 17 references against 16, with the balance shifted from 8 self-citations to 5.
R1-Rev2-Q27 — Reconsider “validated” and “largest and most complete”
“The conclusion describes the resource as a validated dataset (…). These statements should be reconsidered after the methodological issues described above are addressed. Similarly, ‘largest and most complete (…)’ should either be objectively demonstrated or softened.”
Response. Reconsidered after addressing the issues, as the Reviewer instructs, and the two phrases resolve differently.
“Validated” is retained, on the basis set out under Q1: agreement between sources is now measured over 559 race-years and 4,732,776 discipline pairs rather than six races, alongside completeness, internal consistency and physiological plausibility computed over the whole dataset. We have also stated what the agreement does not establish (Q13), so the word is bounded rather than merely asserted.
“Most complete” is removed. “Largest openly available” is retained, with its comparator stated in the sentence.
The Conclusions now also report where the resource is weaker, in the same paragraph as its strengths: agreement is lower in the earliest seasons, transition times are absent for roughly one record in six, and the fields carried only by the official platform are unavailable for a quarter of the dataset.
Corrections we identified ourselves
Working through the comments surfaced errors that neither reviewer raised. Under Open Review we would rather report them than leave them to be found. Two are described above, in the responses to Q18 and Q20, because they share a cause with the points the reviewers made. Two more are recorded here.
The country count was not a count of countries. The submitted Results stated that athletes from 490 countries are represented. The country field is free text, and its distinct values include dropdown placeholders, US state abbreviations, bare initials, code fragments and misspellings; there are approximately 195 countries in the world, and any figure near 490 should have been implausible on its face. The claim is replaced by one computed from the ISO 3166-1 field: 251 distinct country codes, present for 74.2% of records. The leading-country shares were verified and are correct, and are retained from the free-text field because it covers 99.6% of records — the United States (937,613; 34.6%), the United Kingdom (183,093; 6.8%) and Australia (152,980; 5.7%). Their rank order was wrong in the submitted text and is corrected. The condition of the country field is now documented in Table 1 and in the reuse guidance, since it is exactly the kind of trap Rev1-Q7 asks us to warn reusers about.
Fifty-five records carried an upstream redaction we had not propagated. Fifty-five records hold a redaction marker in both the name and country fields. These come from the supplementary source, which honours suppression requests; the marker records a privacy decision taken by an athlete. Our deposit removes names for everyone, so the redacted name was invisible inside the general de-identification — but the redacted country would have survived into the deposited file, together with the event, age group, bib number and split times of people who had asked not to be identifiable. Those 55 records have been withdrawn from the deposit. The consequence is stated in the manuscript rather than smoothed over: the deposited file contains 2,706,867 records while the dataset described in this article comprises 2,706,922, and Table 1 marks which fields reach the deposit.
A new version of the Zenodo deposit (version 3.0.0, https://doi.org/10.5281/zenodo.22097794) has been published carrying the corrected file, the de-identified wording, and the full eight-author list. The Data Availability Statement now cites the concept DOI, which always resolves to the current version.
Withdrawing the records from the newest version was not by itself sufficient, and we mention this because it is the kind of half-measure that can pass for a fix. Zenodo versions are permanent — a published DOI must keep resolving and cannot be deleted — so the earlier version continued to distribute the file those 55 records were in, which is the version the submitted manuscript cited. Access to that file has now been closed. The earlier DOIs still resolve, and their record pages state that the version is superseded, why the file was withdrawn, and where the current data are.
Summary of changes
|
Comments addressed |
43 (1 editorial, 10 Reviewer 1, 27 Reviewer 2, 5 carried from the pre-check round) |
|
Comments rejected |
0 |
|
New analysis scripts |
11 |
|
Supporting scripts (figure and table generation, numeric audit) |
3 |
|
Figures regenerated |
4 |
|
Figures added |
1 |
|
Tables rebuilt from generated output |
2 |
|
References added / removed |
4 / 3 |
|
Self-citation rate |
50% → 29.4% |
|
Numeric claims verified against persisted results |
100 of 100 |
We thank the Editor and both Reviewers again. The manuscript is materially more accurate than the version they received, and in several places that is a direct consequence of questions they asked.
Round 2
Reviewer 2 Report
Comments and Suggestions for AuthorsThank you for the substantial revision of the manuscript.
The authors have satisfactorily addressed the major methodological concerns raised in my previous review.
I only have two minor remaining comments.
In the opening paragraph of the Discussion please replace "two independent sources" with two separate data sources or with something like this, as the manuscript now correctly clarifies that the sources are not independent measurements.
Please check the order/numbering of Figures 4 and 5, as Figure 5 currently seems to appear before Figure 4.
Thank you very much.
Author Response
Response to Reviewers — IRONMAN® Data Paper R2
Manuscript: “A Validated, Population-Scale Dataset of 2.7 Million IRONMAN® Triathlon Records with Separated Transition Times (2002–2026)” Manuscript ID: sci-4491984 Journal: Sci (MDPI) R2 decision received: 2026-08-31 R2 response date: 2026-08-31
Note on Open Review: this manuscript is under Open Review, so this response is published alongside the paper. It is written to be read by anyone, not only by the reviewer.
We thank Reviewer 2 for reading the revision closely enough to find the sentence we forgot to bring with us, and for noticing an ordering defect that had been in the manuscript since the first submission.
Both comments are accepted. Neither is contested. Both go slightly further than the reviewer asked, and we say where and why below.
Reviewer 2
R2-Rev2-Q1 — “two independent sources” in the opening paragraph of the Discussion
“In the opening paragraph of the Discussion please replace ‘two independent sources’ with two separate data sources or with something like this, as the manuscript now correctly clarifies that the sources are not independent measurements.”
Response. Corrected, using the reviewer’s own wording. The Discussion now opens: “…it is drawn from two separate data sources with documented cross-validation…”
The reviewer has caught an inconsistency of our own making. The previous round conceded, in Methods and again in Limitations, that the second source republishes results produced by the same timing operation as the first, so agreement between them evidences faithful transcription rather than independent measurement. That concession is the substance of the revision. We then left the opening sentence of the Discussion asserting the claim we had just given up. A reader arriving at the Discussion first would have taken the stronger claim as the paper’s summary of itself.
We have extended the correction to two further sentences the comment does not name. In both, “independent” was used in a different and defensible sense — independent of the IRONMAN organisation, which is true — but it is the same word doing different work in one document, and after this comment no reader should have to decide which sense is meant:
- Abstract: “an independent aggregator covering event series absent from it” → “a third-party aggregator covering event series absent from it”.
- Methods, Data sources: “CoachCox is an independent statistics aggregator” → “CoachCox is a third-party statistics aggregator”.
The sentence immediately following the second, which draws the distinction the reviewer is pointing at, is unchanged: “Its independence from the official platform is editorial rather than metrological: it collects and republishes results produced by the same timing operation.”
All three changed passages are highlighted in the manuscript file.
R2-Rev2-Q2 — the order and numbering of Figures 4 and 5
“Please check the order/numbering of Figures 4 and 5, as Figure 5 currently seems to appear before Figure 4.”
Response. Corrected, and the correction is larger than the pair named, because the defect is.
The reviewer saw the visible symptom. Underneath it, neither sequence in the manuscript was monotonic. The body cited the figures in the order 1, 4, 2, 5, 3, while the document laid them out 1, 2, 5, 4, 3 — the two disagreed with each other as well as with the numbering. The specific cause of what the reviewer saw is that the temporal-trends figure was cited early, in the second paragraph of the Results, and placed late, four subsections further down; the sex-participation figure sat in between.
Swapping the two figures the comment names would have removed the symptom and left both underlying problems in place. We have instead renumbered all five figures in order of first citation, so that citation order, numbering and physical position now agree:
|
Previous |
Now |
Figure |
|
1 |
1 |
Dataset composition by data source and race type |
|
4 |
2 |
Temporal trends in participation and performance |
|
2 |
3 |
Split time distributions for full-distance finishers |
|
5 |
4 |
Participation by sex across the covered period |
|
3 |
5 |
Transition time coverage |
Each figure has also been placed at the paragraph that first cites it.
Two things follow from this that the reviewer and any later reader should know:
- No image changed. The five image files were renamed, not regenerated; each carries the same checksum it had before. Every caption is the text it had before, attached to the same figure. Only the numbers and the order changed.
- The R1 response letter, published with these reviews, uses the previous numbering. Where it says Figure 5 was added to answer the participation-by-sex comment, that figure is now Figure 4; where it discusses the removal of a redundant panel from Figure 4, that figure is now Figure 2. We have not rewritten the earlier letter, because it is the record of what was said at the time.
This ordering defect predates the first submission. It was found at the end of the previous round, after the response letter had gone to the editor, and deferred rather than fixed then — renumbering at that moment would have desynchronised the manuscript from a letter already in the reviewers’ hands. This is the round in which it belonged, and the reviewer arrived at it independently.
Reviewer 1
No comments were returned in this round.
Changes not requested by a reviewer
One, and it is administrative: the correspondence address for Pantelis Theodoros Nikolaidis in the author list has been changed to pademil@hotmail.com. Mail sent to the address carried in the previous version — including the journal’s own correspondence to the author list — has been bouncing since 19 August.
No number, analysis, table or reference changed in this round.
Author Response File:
Author Response.docx