Lactate Test Progress Tracking: The 16-Field Log That Makes Trends Readable
Lactate test progress tracking: a 16-field session record, a noise band you can calculate, three trend series, retest triggers and a worked 22-week example.
Track three series per test, not one number: the load at LT1, the load at LT2, and blood lactate at a single reference load you keep identical in every test. The third series moves first, and it is the one most home logs never record. Everything else in a tracking system โ the fields you write down, the retest cadence, the method you use to derive thresholds โ exists to keep those three series comparable across months. Logs usually fail not because the measurements were bad, but because something in the protocol changed between test one and test four and nobody wrote it down.
Three series, not one number
Most training logs store a single figure after each step test: threshold watts, or threshold pace. Then, six weeks later, someone argues about whether five watts means anything. It usually doesn't โ and the reason is geometric rather than physiological.
Reading a threshold means reading a position on the x-axis. Reading the reference stage means reading a value on the y-axis at a fixed x. The vertical gap between two lactate curves at a fixed load is roughly the horizontal shift multiplied by the local slope of the curve, so the same underlying change shows up amplified wherever the curve is steep. Between LT1 and LT2 the curve is already rising steeply, at a load you can still repeat identically in every test. That is where you want your reference stage.
Reference-stage selection rule: from your baseline test, pick the stage where lactate sat somewhere between about 2 and 4 mmol/L, and keep that exact load in every subsequent test even after it stops feeling hard. If your ladder is 120โ300 W in 20-watt steps, the reference stage is a fixed number like 220 W, not "the fourth stage" โ stage numbering drifts as soon as you change the start load.
Record heart rate at that same reference stage as a fourth, weaker series. It is noisier than lactate and sensitive to sleep, caffeine and room temperature, but a heart rate that drifts in the opposite direction to lactate is a useful prompt to check your context fields before you believe anything.
What to record: the 16-field session record
Seven of these fields are locked. If a locked field changes, you have not made a new entry in the same series โ you have started a second series that cannot be compared with the first, no matter how tempting the numbers look side by side.
| # | Field | Example entry | Status | |---|---|---|---| | 1 | Modality and ergometer | Wattbike Atom, unit #2 | Locked | | 2 | Start load | 120 W | Locked | | 3 | Step size | 20 W | Locked | | 4 | Stage duration | 3:00 | Locked | | 5 | Sampling site | Right earlobe | Locked | | 6 | Sampling moment in stage | Final 20 s, pedalling | Locked | | 7 | Analyser model + strip lot | Lactate Plus, lot 4471B | Locked | | 8 | Time of day | 17:30 | Context | | 9 | Training in the previous 48 h | Rest day + 60 min easy | Context | | 10 | Carbohydrate intake in the 3 h before | ~60 g, 2 h before | Context | | 11 | Venue, temperature, humidity | Garage, 21 ยฐC, 55 % | Context | | 12 | Body mass on test morning | 74.2 kg | Context | | 13 | Raw table: load, HR, lactate, RPE per stage | attached CSV | Raw | | 14 | Resting/baseline lactate | 1.1 mmol/L | Raw | | 15 | Final stage + completed in full? | 300 W, aborted at 2:10 | Raw | | 16 | Threshold method + tool version | Dmax, LactateThreshold, Aug 2026 | Derived |
Fields 13 to 16 are the reason to keep raw data rather than only the pretty summary. A derivation method can be re-applied to a two-year-old test; the physiology of that afternoon cannot be re-measured. If you keep the raw stage table, changing your mind about Dmax versus a fixed 4 mmol/L definition costs you an afternoon of recalculation instead of an entire archive.
One field people skip and later regret: the strip lot number. It is printed on the vial, it takes four seconds to copy, and it is the first thing to check when a curve sits implausibly high or low with no other explanation.
Compute your noise band before you read any trend
Before you interpret a change, work out the smallest change your setup can actually resolve. The arithmetic takes two minutes.
1. Find the coefficient of variation (CV) in your analyser's package insert. Check the document that came with your meter and strips rather than assuming a figure from a forum โ manufacturers state precision at specific concentration levels, and the value differs between systems and between within-run and between-run conditions. 2. Convert the percentage into mmol/L at the level you care about. Worked example with an assumed 4 %: at a reference-stage value of 3.4 mmol/L, that is ยฑ0.14 mmol/L on a single sample. 3. Combine the two tests you are comparing. Two independent measurements, each carrying that uncertainty, give roughly 0.14 ร โ2 โ 0.19 mmol/L for the difference between them. 4. Treat that number as a floor, not a total. It describes the analyser and the strip. Day-to-day biological variation โ glycogen state, the previous evening's session, hydration, ambient heat โ rides on top of it and is exactly what fields 8 to 12 are there to document.
Because the analytical band is only a floor, we apply a deliberately blunt house rule on top of it: a difference below 0.3 mmol/L at the fixed reference load, or below one stage increment in threshold load, is recorded as unresolved rather than as progress or regression. Unresolved is a legitimate result. Writing it down is more honest than rounding a 4-watt difference into a training-block verdict.
You can also measure your own repeatability instead of borrowing a number. On one stage per test, run two strips from the same drop of blood. That costs one extra strip per session and, after four or five tests, gives you a spread from your own meter, your own hands and your own sampling technique. For the two-test comparison maths in detail, see How to Compare Lactate Test Results Over Time Without False Precision.
Worked example: one athlete, three tests, 22 weeks
The dataset below is fictional and constructed to show how the three series behave differently. Same ergometer, same 120 W start, same 20-watt steps, same 3-minute stages, same earlobe sampling in the final 20 seconds.
| Series | Test A (week 0) | Test B (week 11) | Test C (week 22) | |---|---|---|---| | Lactate at 220 W reference stage | 3.4 mmol/L | 2.9 mmol/L | 2.6 mmol/L | | LT1 load | 196 W | 204 W | 212 W | | LT2 load (Dmax) | 268 W | 272 W | 281 W | | Heart rate at 220 W | 158 bpm | 155 bpm | 154 bpm | | Body mass | 74.2 kg | 73.5 kg | 72.8 kg | | LT2 in W/kg | 3.61 | 3.70 | 3.86 | | Baseline lactate | 1.1 mmol/L | 0.9 mmol/L | 1.0 mmol/L |
Now read it against the band. The reference stage fell by 0.8 mmol/L across the 22 weeks โ more than four times the ยฑ0.19 mmol/L analytical band and well outside the 0.3 mmol/L house floor. That series is resolvable.
The headline number is not. LT2 gained 13 watts in total, which is less than a single 20-watt stage increment. A curve fit will happily report LT2 to the watt, but its confidence comes from stages spaced 20 watts apart, so sub-increment movement should stay provisional. What rescues it here is direction: 268 โ 272 โ 281 is monotone across three tests, and a direction repeated three times is a stronger argument than any single jump.
The sentence worth pinning to the log: in 22 weeks the threshold number moved less than one stage step, while the reference stage moved more than four times its own noise band.
The W/kg row is the trap. It rose from 3.61 to 3.86, an eye-catching movement โ but part of that is the 1.4 kg lost in the denominator. Keep mass in a separate column and never let a normalised value be the only thing you archive.
Why the reference stage moves first
At a fixed load, you are sampling the curve where its slope is already steep, so a modest horizontal shift produces a comparatively large vertical drop. At the threshold, you are asking a fitted model to locate a breakpoint using stages 20 watts apart. The first is a measurement; the second is an estimate built on top of a measurement, and estimates inherit every uncertainty underneath them.
There is a practical consequence for retest timing. If a block produces a change too small to shift a threshold past one stage increment, the reference stage may still show it โ which is why a test that returns "LT2 unchanged" is not automatically a wasted afternoon. Read the y-axis before you conclude nothing happened.
What the log tells you is which direction the data moved. What to do about it is a training decision for you and your coach, and any health question the numbers raise belongs to a qualified professional.
Protocol resolution: why 20-watt steps cap what you can detect
The obvious fix for coarse resolution is finer steps. Run the arithmetic before you commit to it.
| Ladder | Stages from 120 to 300 W | Time at 3-min stages | Strips per test (incl. baseline) | |---|---|---|---| | 20 W steps | 10 | 30 min | 11 | | 10 W steps | 19 | 57 min | 20 |
Halving the step size roughly doubles the test duration and the strip count. A 57-minute incremental test is a different physiological demand from a 30-minute one โ accumulated fatigue, fuelling and thermal load all change โ so the finer ladder does not simply give you a higher-resolution version of the same test. It gives you a different test, which starts a new series.
The cheaper lever is narrowing the range rather than the steps. Once you know roughly where the interesting part of the curve sits, start higher and keep the same 20-watt spacing over fewer stages. You keep the step geometry, shorten the test and can afford a duplicate strip.
Budget note for a season: 10 stages plus baseline is 11 strips, plus one duplicate makes 12. Three tests a year is roughly 36 strips before any repeat or failed sample โ worth checking against the pack size and expiry date you actually own.
When the number moves because the method changed
Threshold values are definitions applied to data, not readings taken from a body. Dmax, a fixed 4 mmol/L cut-off, and baseline-plus-a-fixed-increment will each return a different load from the same seven rows, and the gaps between them are frequently larger than a season of training.
Three rules keep this from contaminating a log:
- Store the method name and tool version with every result (field 16). "LT2 = 268 W" is an incomplete record.
- When you change methods, recompute every archived test with the new method before you compare anything. This is only possible if you kept raw stage tables.
- Keep one method as the tracking series and treat the others as cross-checks that live in their own columns.
If you are choosing or switching tools, the practical requirement for tracking is narrow: the calculator must accept your raw step-test rows, name the LT1 and LT2 method it applied, and let you export the data again. A tool that shows a lactate curve but will not give the underlying numbers back is a dead end for a multi-year log. Calculate lactate curve online: step-test input and threshold output walks through what an LT1/LT2 calculation needs as input and what it returns.
Test cadence: what should trigger the next test
These are comparability rules for the log, not training advice.
| Situation | Rule we apply | |---|---| | A training block of 8+ weeks with stable structure has finished | Test โ the block is the unit the series compares | | You changed ergometer, treadmill or sport | Test if you want, but open a new series; the old one ends | | Returning from illness or a layoff | Wait until you can complete the full stage ladder; an aborted ladder truncates the curve and the fit | | A single session felt unusually good or bad | Do not test. One session is not a hypothesis a step test can answer | | Race week | A test in a tapered state measures a different condition from the block tests it would sit next to. If you run it anyway, flag it in the log so it is never read as part of the trend | | Heat wave, or you moved to altitude | Record the venue conditions and expect the comparison to be compromised; do not silently fold it into the series |
The most common cadence failure is testing too often to confirm a feeling. Tests spaced four weeks apart mostly sample the noise band; tests spaced around a training block sample the block.
Where wearables and continuous monitors fit in the log
Continuous and wearable lactate devices are an active development area โ a 2024 narrative review by Yang and colleagues surveys devices engineered to monitor lactate in sweat during sport (PMC11025537). For a progress log, the operational question is narrower than the technology debate.
Sweat and capillary blood are different sample matrices. A series built from capillary blood samples and a series produced by a sweat-based wearable are two logs, not one, and merging them retroactively destroys both. Before combining anything, read the manufacturer's own validation documentation and check what it was compared against, in which sport, and over what intensity range.
Where a continuous device is genuinely useful in a tracking context is between tests: it can produce a dense within-session record where a step test gives you seven points a few months apart. Keep it in its own column, label it by device and firmware version, and hold it to the same locked-field discipline you apply to the blood series.
Five ways a multi-year log goes bad
1. The strip lot changed mid-season and nobody logged it. The tell is a whole curve that shifts up or down while the shape and the heart-rate data stay put. 2. Sampling moved from earlobe to fingertip. Usually because the earlobe stopped cooperating on a cold morning. The tell is a single test that breaks an otherwise consistent direction, with no matching change in the context fields. 3. Stage duration drifted from 3:00 to 4:00 โ often because new software defaulted to it. The tell is the last stages sitting higher than the trend while the early stages look normal. 4. Raw values were overwritten with "corrected" ones. Now nothing can be recomputed. Corrections belong in a new column with a note, never on top of the original row. 5. Only the newest test was recalculated with a new method. This produces the most convincing fake progress there is, because the jump lands exactly where you were hoping to see one.
The minimum sheet you can start this week
One row per stage, one sheet per test, one folder per athlete:
`test_date | stage_no | load | duration | HR | lactate | RPE | notes`
Plus a header block carrying fields 1 to 12 and 16, and a single line at the top recording the reference load you chose from the baseline test.
If you have already run tests and did not record the locked fields, do not throw the old data out. Mark those tests as "series 0", write down whatever you can reconstruct with confidence, and start the disciplined series with the next test. A clean two-point series beats a dirty six-point one. If you are still assembling the measurement side of this, Self test lactate analysis online: measure and interpret results covers the sampling workflow the log depends on.
Questions people ask about lactate test progress tracking
How do I read lactate test results across several tests? Read them as three parallel series โ reference-stage lactate, LT1 load, LT2 load โ each against the noise band you calculated, with the context fields open next to them. Read the shape of the curve alongside the numbers. What the values mean for your training is a decision for you and your coach, not something the spreadsheet decides.
How long does a lactate step test take? It depends on stage duration and how many stages you run. UC Davis Sports Medicine describes a lactate profile protocol in which the load is increased gradually every 3โ5 minutes until one stage above the point where lactate reaches 4 mmol/L (UC Davis Sports Medicine). With 3-minute stages and a ten-stage ladder, plan around 30 minutes of stages plus warm-up, sampling handling and cool-down.
What lactate level is concerning? That question comes from clinical medicine, where blood lactate is measured for reasons unrelated to a step test and interpreted against clinical reference ranges by clinicians. Exercise-test values from a step test are not interchangeable with those measurements, and nothing in a training log is a substitute for medical assessment. If you have a health concern, take it to a physician.
Should I retest with a different protocol if the current one feels too easy at the start? You can, and sometimes should โ but it ends the current series. Note the change, keep the old series visible, and expect to run two tests on the new protocol before you have a direction again.
What the log cannot tell you
It cannot tell you why anything moved. A reference stage that drops 0.5 mmol/L looks the same whether the cause was 11 weeks of training, an unusually well-fuelled test day, or a warmer garage. The context fields narrow the field of candidates; they do not close it. And a log built on locked fields still cannot answer the question most people actually bring to it โ whether the training that produced the change is the training to keep doing. That question needs a coach, a race result, or another season of data. The log's job is only to make sure the numbers you argue about are worth arguing about.
Read next: Calculate lactate curve online ยท Compare lactate test results over time.