Data Collection
What is Data Collection?
Data collection is the process of gathering information to answer a question or test a hypothesis. Before collecting any data, you need to decide:
- What question are you trying to answer?
- What type of data do you need?
- Who or what will you collect data from?
- How will you collect it?
Good data collection leads to reliable conclusions. Poor data collection โ even with excellent analysis โ can produce misleading results.
Types of Data
Qualitative (Categorical) Data
Qualitative data describes categories or qualities that cannot be measured numerically.
Examples: eye colour, favourite film genre, country of birth, type of transport used.
This data is usually summarised using tallies and represented with bar charts or pie charts.
Quantitative (Numerical) Data
Quantitative data is numerical and can be measured or counted.
It comes in two subtypes:
Discrete data โ takes only specific, separate values (usually whole numbers). You can count it.
- Examples: number of siblings, shoe size, goals scored in a match
Continuous data โ can take any value within a range. You measure it.
- Examples: height, weight, temperature, time, distance
The distinction matters because continuous data is displayed with histograms, while discrete data typically uses bar charts.
Primary vs Secondary Data
| Type | Definition | Examples |
|---|---|---|
| Primary | Collected by you directly for your specific investigation | Surveys, experiments, observations |
| Secondary | Collected by someone else and used by you | Census records, newspaper data, published research |
Primary data is tailored to your exact question but takes time and effort to gather.
Secondary data is quicker to access but may not perfectly match your needs and you cannot control its quality.
Sampling Methods
When a population is too large to collect data from everyone, you take a sample โ a smaller, representative group.
Random Sampling โ every member of the population has an equal chance of being selected.
- Advantage: unbiased. Disadvantage: requires a full list of the population.
Systematic Sampling โ select every nth member from a list.
- Example: from a list of 100, select every 5th person (gives a sample of 20).
- Advantage: simple and spread out. Disadvantage: can miss patterns that repeat at the same interval.
Stratified Sampling โ the population is divided into subgroups (strata) and members are selected from each in proportion to the subgroup size.
- Example: if 60% of students are female, 60% of the sample should be female.
- Advantage: representative of all subgroups. Disadvantage: requires detailed population information.
Quota Sampling โ similar to stratified but not random within subgroups; the researcher fills quotas until the target number is reached.
Convenience Sampling โ use whoever is easiest to reach.
- Advantage: fast and cheap. Disadvantage: highly likely to be biased.
Census vs Sample
| Census | Sample | |
|---|---|---|
| Definition | Data from the entire population | Data from a subset of the population |
| Accuracy | Completely accurate (no sampling error) | Subject to sampling error |
| Cost | Expensive and time-consuming | Cheaper and quicker |
| Practicality | Impractical for large populations | Practical for most situations |
A national census (such as the UK census every 10 years) attempts to collect data from every household in the country.
Good Questionnaire Design
A well-designed questionnaire produces reliable, unbiased data. Key principles:
- Clear and unambiguous โ every respondent should interpret each question the same way.
- One idea per question โ do not ask two things at once ("Do you eat fruit and vegetables daily?").
- Avoid leading questions โ do not steer respondents towards a particular answer. ("Do you agree that cycling is healthy?" is biased.)
- Response options must be exhaustive and mutually exclusive โ every possible answer should fit exactly one option.
- Avoid sensitive questions at the start โ build rapport before asking age, income, or personal topics.
- Pilot the questionnaire โ test it with a small group to find unclear or problematic questions before the main survey.
Example of a poor question: "How often do you sometimes exercise?"
- "Sometimes" is vague; response options may overlap.
Improved version: "How many times per week do you exercise? (0 / 1-2 / 3-4 / 5 or more)"
Key Terms
| Term | Meaning |
|---|---|
| Population | The entire group being studied |
| Sample | A subset of the population |
| Bias | Systematic favouritism that distorts results |
| Hypothesis | A statement to be tested with data |
| Pilot study | A small-scale trial run of the investigation |
| Leading question | A question that pushes the respondent toward a particular answer |
Common Mistakes
- Confusing discrete and continuous data โ the number of people in a room is discrete (you cannot have 3.5 people); their heights are continuous.
- Using convenience sampling and claiming it is random โ a random sample requires every member to have an equal chance, not just those nearby.
- Overlapping response categories โ if "20-30" and "30-40" are both options, what does a person aged 30 choose?
- Forgetting a "none/other" option โ always provide a catch-all for responses that do not fit neatly into any category.
- Not piloting the questionnaire โ unclear questions only become obvious when real respondents try to answer them.
Tips and Tricks
- Always state your hypothesis or research question clearly before choosing a data collection method.
- Match the sampling method to your resources: random sampling is most rigorous but requires a full population list.
- For continuous data, use class intervals (e.g. 10 โค x < 20) to group values โ make sure intervals do not overlap and together cover all possible values.
- A good sample size is large enough to be representative but small enough to be manageable โ typically at least 30 for basic statistical work.