In simple terms
A friendly intro before the formal notes — no formulas yet.
Testing Athletes: Are We Measuring the Right Thing, Right?
In sports science, we need reliable tools to measure an athlete's fitness. This involves choosing the correct test for the job and understanding how to analyse the results fairly and accurately.
Think of baking a cake. To get a perfect Victoria sponge, you need a specific recipe for that cake, not one for a loaf of bread. This is 'validity' – using the right tool for the job. You also need to be able to follow that recipe exactly the same way every time to get a consistent result. This is 'reliability' – the consistency of your measurement.
- 1
First, select a test that is both valid (it measures the specific fitness component you're interested in) and reliable (it produces consistent results).
- 2
Next, standardise the testing procedure. This means controlling variables like the warm-up, equipment, time of day, and environment to ensure a fair comparison.
- 3
Then, conduct the test and collect the performance data accurately. Ensure instructions are clear and scoring is objective.
- 4
Finally, analyse the data using statistics. Calculate the mean to find the average, standard deviation to see the spread of scores, and use a t-test to compare different groups or conditions.
Explore the concept
Use the live diagram, PhET or GeoGebra sim, and synced steps — play it, drag controls, or tap a step.
Step 1
First, select a test that is both valid (it measures the specific fitness component you're interested in) and reliable (it produces consistent results).
1 more simulation for this topic — run them in the Simulations section below
Simulations
Every simulation here runs the real model — try the steps on a card, then check what you see against the notes.
1 simulation
- GeoGebraCoreIB 2.1
Visual Demo of Standard Deviation
Twenty lettered points are test scores on a number line; drag any of them left or right and the average and standard deviation update.
Try this
- Read the starting average and standard deviation for the squad's scores.
- Drag one athlete's score far to the right and compare how much the standard deviation jumps with how little the average moves.
- Bunch every score close together, then spread them out, and compare the two standard deviations.
- Each time, divide the standard deviation by the average to get the coefficient of variation.
Look for Standard deviation measures how far scores sit from the mean: a consistent squad has a small value, and a single outlier inflates it.
Terry Lee Lindenmuth · GeoGebra · GeoGebra Terms of Service
Key formulas
Tap any symbol to reveal exactly what it means and its units.
Tap a symbol — great for exam definitions
Full topic notes
Formal explanation with the rigour you need for the exam.
Core Principles of Measurement
To ensure that the data we collect is meaningful, our tests must adhere to several key principles. The two most important are validity and reliability. Validity ensures we are measuring the correct component of fitness, while reliability ensures our measurements are consistent over time. Imagine trying to measure an athlete's aerobic endurance using a vertical jump test; this would not be a valid measure. Similarly, if a stopwatch gives wildly different times for the same performance, it is not reliable.
Validity: The test measures what it is supposed to measure.
Reliability: The test produces consistent results upon repetition.
Accuracy: The degree to which a measurement corresponds to the true value. This is related to the quality of the measuring instruments.
Objectivity: The degree to which different testers will produce the same result for the same subject. This is a component of reliability.
Designing a Fitness Test Protocol
When designing a fitness test, it is crucial to standardise the procedures to ensure reliability and allow for valid comparisons. This involves controlling as many variables as possible. Before any testing, participants should complete a Physical Activity Readiness Questionnaire (PAR-Q) and provide informed consent to ensure safety and ethical practice.
Specificity: The test must be relevant to the sport or activity (e.g., testing agility for a tennis player).
Standardised Warm-up: All participants should perform the same warm-up.
Consistent Order of Tests: If conducting a battery of tests, the order should be the same for everyone to control for fatigue.
Controlled Environment: Factors like temperature, humidity, and surface should be kept constant.
Calibrated Equipment: All measurement tools (scales, stopwatches, etc.) must be checked for accuracy.
In exam questions, when asked to design a test, don't just name it. You must justify your choice by linking it to a specific component of fitness and explaining the protocol, including standardisation procedures, to ensure validity and reliability. For example, 'To measure aerobic endurance, the Multistage Fitness Test would be used. It is a valid test of VO2 max. To ensure reliability, the test would be conducted on the same indoor surface, with the same audio track, and a standardised 10-minute warm-up.'
Analysing Data: Mean and Standard Deviation
Once data is collected, we use statistics to describe and interpret it. The most common measure of central tendency is the mean, or average, of the scores. This gives us a single value that represents the typical performance of the group.
Mean () =
However, the mean doesn't tell us about the consistency of the scores. For that, we use the standard deviation (SD). The SD tells us how spread out the scores are from the mean. A small SD means the scores are tightly clustered around the average, indicating consistent performance. A large SD means the scores are widely spread, indicating inconsistency.
Standard Deviation (s) =
Comparing Data: The t-test
Often, we want to know if a difference between two groups is 'real' or just due to random chance. For example, did a new training programme significantly improve performance compared to a control group? To answer this, we use an inferential statistical test, such as the t-test. The t-test compares the means of two groups and tells us the probability (p-value) that the observed difference happened by chance.
A t-test is used to compare the means of two groups.
The null hypothesis () states there is no significant difference between the groups.
The alternative hypothesis () states there is a significant difference.
We compare our calculated t-value to a critical value from a table, or we look at the p-value.
If p < 0.05, we reject the null hypothesis and conclude there is a statistically significant difference.
Worked examples
See the formulas applied — reveal one step at a time, like the exam.
A group of five sprinters recorded the following times for a 30-metre sprint: 4.2s, 4.5s, 4.1s, 4.3s, 4.4s. Calculate the mean and standard deviation for this data set. (You may use a calculator's statistical function for SD).
- 1
Calculate the Mean ():
A researcher investigates the effect of a 6-week plyometric training programme on the vertical jump height of basketball players. Group A (n=15) followed the programme, while Group B (n=15) was a control group. The mean improvement for Group A was 5.2 cm, and for Group B was 1.1 cm. A t-test was performed on the data, yielding a p-value of p = 0.03. Using a significance level of p < 0.05, interpret these results.
- 1
State the significance level: The chosen level for significance is p < 0.05. [1 mark]
How it all connects
The big idea sits in the middle — tap a linked idea to explore the link.
Tap a linked idea to see how it connects back to the main topic — that connection is what examiners reward.
Glossary
Key terms for this topic — skim now; the Check step will test them.
- validity in fitness
Validity refers to whether a test actually measures what it claims to measure. For example, using a 1-repetition maximum (1RM) bench press to test maximal strength is valid.
- reliability in fitness
Reliability is the degree to which a test is consistent and stable in measuring what it is intended to measure. If an athlete performs the same test on two separate occasions, the results should be approximately the same.
- Standard Deviation (SD)
A measure of the spread or dispersion of a set of data from its mean. A low SD indicates that the data points tend to be close to the mean, whereas a high SD indicates that the data points are spread out over a wider range.
- purpose of a t-test
A statistical test used to determine if there is a significant difference between the means of two groups. For example, comparing the effectiveness of two different training programmes on vertical jump height.
- PAR-Q
A Physical Activity Readiness Questionnaire. It is a self-screening tool used to identify any potential health risks before an individual starts a new physical activity programme or fitness test.
- objectivity in testing
A measure of how free a test is from the biases of the tester. An objective test will produce the same result regardless of who is administering it. For example, using timing gates for a sprint is more objective than using a handheld stopwatch.
Quick check
Write your answer first, then compare it with the model one — the gap is what you would have lost.
Teach it back
If you can explain it simply, you own it — gaps here are marks you’d lose.
Teach it back
Explain this topic as if teaching a friend. We name the gaps an examiner would still dock.
Revision flashcards
Guess first, then flip — retrieval beats re-reading.
Key takeaways
Review these before you close the topic — retrieval beats re-reading.
Validity: The test measures what it is supposed to measure.
Reliability: The test produces consistent results upon repetition.
Accuracy: The degree to which a measurement corresponds to the true value. This is related to the quality of the measuring instruments.
Objectivity: The degree to which different testers will produce the same result for the same subject. This is a component of reliability.
Practice — then mark it
The whole point: a real Cambridge question, marked mark-by-mark.
Test Your Knowledge on Measurement and Evaluation
Test Your Knowledge on Measurement and Evaluation
Extra simulations & links
PhET, GeoGebra and other curated tools — open in a new tab.
Frequently asked
Checkpoint
One marked question is worth ten re-reads — close the loop before you move on.
Reading it isn’t knowing it — prove it.
Before you move on: do Test Your Knowledge on Measurement and Evaluation on paper, snap a photo, and get examiner-style feedback on exactly where you win and lose marks.
Discuss Measurement and evaluation of performance
Ask, share and discuss with other Sports, Exercise and Health Science SL students
No posts yet — be the first to start the conversation.