Overview
This unit, Methods of Psychology, introduces the scientific ways psychologists study behaviour and mental processes. It covers how questions are formed, how hypotheses are tested, and the main research methods used in psychology such as experiments, correlational studies, surveys, observations and case studies. The unit also explains measurement, sampling, reliability, validity, and the ethical standards required when working with people. Learning these methods helps students understand how psychological knowledge is produced, how to evaluate claims, and how to design their own simple studies. These skills matter because they form the foundation for later work in psychological assessment, counselling, education, health, and research. Students will learn to distinguish between cause and correlation, to recognise bias, and to analyse data responsibly. The unit emphasises practical competence: planning studies, choosing appropriate methods, collecting and recording data, and presenting findings clearly. Through examples and practice questions, students gain the ability to read research critically and to apply scientific thinking to everyday observations about human behaviour.
Learning Objectives
- Define key research terms used in psychology and explain their importance.
- Differentiate between major research methods such as experiments, surveys, observations and case studies.
- Formulate testable hypotheses and describe how variables are identified and operationalised.
- Explain sampling techniques and justify their use in different research contexts.
- Assess the reliability and validity of psychological measures and studies.
- Apply ethical principles to plan research involving human participants.
- Collect, record and present basic quantitative and qualitative data accurately.
- Interpret simple statistical results and draw appropriate conclusions from data.
Topics in this chapter
17 topics · tap a topic title to jump straight to it.
Introduction to Methods in Psychology
What the methods area covers
Psychology seeks to explain behaviour and mental processes in ways that are reliable and generalisable. To do this, psychologists use systematic methods that help them move from curiosity or anecdote to evidence. This topic introduces why methods matter, what makes a method scientific, and the difference between everyday explanations and research-based explanations.
Science and psychology
Scientific study in psychology means observations are planned, procedures are repeatable, and conclusions are based on collected evidence. Psychology uses both quantitative and qualitative approaches. Quantitative methods measure variables numerically, while qualitative methods describe experiences, meanings and patterns without reducing them to numbers.
Key stages of research
Typical research follows a series of steps: asking a clear question, reviewing what is already known, forming hypotheses, choosing suitable methods, collecting data, analysing findings, and drawing conclusions. Each stage requires careful decisions: an unclear question leads to vague methods; poor measurement leads to misleading results. Students learn to think through each stage and to plan studies so they will produce meaningful, interpretable evidence.
Why methods matter for students
Understanding methods helps students evaluate news, advertisements and claims about behaviour. It supports critical thinking: is a claim based on solid evidence? Could alternative explanations account for the same facts? Students also learn how to design simple classroom investigations, collect and interpret data, and communicate findings clearly. These abilities are useful beyond psychology—for sciences, social studies and everyday reasoning.
Mixed methods and triangulation
No single method answers all questions. Using multiple methods—called triangulation—helps build stronger conclusions. For example, combining survey data with observations and a small experiment can show that a pattern exists, how it looks in the real world, and whether a causal mechanism is plausible. Triangulation compensates for the weaknesses of any single approach.
Limitations and responsible use
Every method has limits: experiments can be artificial, surveys rely on self-report, and observations can be biased. Good research acknowledges limitations and avoids overstating results. Ethical and transparent reporting, replicability, and cautious generalisation are key principles. For students, learning methods is about gaining tools to ask better questions and to be honest about what evidence does and does not show.
- Comparing everyday intuition ('people learn better when rewarded') with experimental evidence that tests learning under controlled reward and no-reward conditions.
- A classroom activity where pupils notice how many classmates prefer a certain study time, then discuss whether this sample represents the whole school.
- Examining a newspaper claim about a new therapy and tracing whether the claim cites scientific studies or only testimonials.
- Hypothesis: A clear, testable statement predicting the relationship between variables.
- Operational definition: A statement of how a variable will be measured or manipulated.
The Scientific Method and Hypothesis Formation
Purpose of the scientific method
The scientific method provides a structured process to investigate questions about behaviour so that findings are systematic, transparent and open to verification. It reduces bias by relying on planned observation, repeatable procedures and logical inference from evidence.
From question to hypothesis
Research begins with a clear, focused question that is observable and measurable. After reviewing previous work and theory, the researcher forms a hypothesis: a specific, testable prediction. A good hypothesis states the expected relationship between variables, and it can be confirmed or disconfirmed by data. Writing a hypothesis forces clarity about what should be measured and how.
Types of hypotheses
There are several forms: directional (predicts the direction of an effect), non-directional (predicts an effect but not direction), and the null hypothesis (which predicts no effect and serves as the default in statistical testing). Researchers often set up both a null and an alternative hypothesis so that statistical tests can evaluate the evidence against the null.
Operationalisation
Operationalisation specifies exactly how abstract concepts will be measured or manipulated. For example, 'stress' can be defined as a score on a validated stress questionnaire, cortisol level in saliva, or number of stressful events reported. Clear operational definitions allow other researchers to replicate the study and help ensure that everyone understands what the variables represent.
Variables and control
Identifying independent, dependent and control variables helps design the study. The independent variable is the presumed cause, the dependent variable is the effect measured, and control variables are factors held constant. Recognising potential confounders at this stage allows the researcher to plan ways to reduce their influence, such as randomisation, matching, or statistical control.
Testing and falsifiability
A scientific hypothesis must be falsifiable—there must be a possible observation that would contradict it. Designing tests that could disprove the hypothesis is central to scientific thinking because it avoids confirming bias and strengthens the credibility of findings that survive rigorous tests.
Replication and cumulative knowledge
Initial studies rarely settle a question alone. Replication—conducting the study again in different conditions and samples—builds confidence. Over time, repeated and varied tests allow science to refine theories, identify boundary conditions, and develop reliable knowledge. Students should value replication and cautious interpretation of single studies.
- Formulating a directional hypothesis: 'Listening to soft music while studying increases test scores.'
- Turning 'happiness' into a measurable variable by using a validated 10-item scale.
- Stating a null hypothesis: 'There will be no difference in memory performance between the two study groups.'
- Independent variable (IV): The variable manipulated.
- Dependent variable (DV): The outcome measured.
- Null hypothesis (H0): No effect or relationship.
- Alternative hypothesis (H1/HA): There is an effect or relationship.
Experimental Method
What is an experiment?
An experiment is a controlled research design where the researcher deliberately manipulates one or more independent variables to observe effects on one or more dependent variables. The central aim is to test causal hypotheses: does X produce change in Y under specified conditions?
Core elements
Key elements are manipulation, control, random assignment and systematic measurement. Manipulation involves changing the independent variable across conditions. Control reduces influence of extraneous factors. Random assignment places participants in conditions by chance, helping ensure groups are comparable. Reliable and valid measures are used to record the dependent variable.
Designs in experiments
Between-subjects designs compare different groups, each experiencing one condition. Within-subjects designs have the same participants undergo all conditions at different times, enhancing statistical power but risking order effects. Counterbalancing (changing the order across participants) reduces such order effects. Matched groups match participants across conditions on key variables to reduce variability when randomisation is difficult.
Controlling confounds
Confounding variables can provide alternative explanations. Techniques to control confounds include randomisation, use of control groups, standardised procedures, and blinding. Single-blind designs conceal condition from participants; double-blind designs also conceal condition from experimenters, reducing expectancy effects. Placebo controls are important in studies of treatments to control for expectancy and other non-specific effects.
Ecological validity and realism
Laboratory experiments maximise control but may reduce realism. Field experiments take place in natural settings to increase ecological validity while maintaining some experimental control. Choosing between lab and field involves trade-offs: lab experiments increase internal validity, while field experiments improve generalisability.
Ethical constraints and practical limits
Not all questions can be addressed experimentally due to ethical concerns (e.g., inducing harm) or practicality (long-term effects). Quasi-experiments provide alternatives when random assignment is impossible, but researchers must carefully address potential confounds and be cautious about causal claims.
Reporting experimental findings
Good reports describe participants, design, procedure, materials and analyses. Including effect sizes, confidence intervals, and replication suggestions helps readers evaluate the strength and applicability of findings. Clear reporting allows others to reproduce and test the study further.
- A lab experiment testing how sleep deprivation (IV: 0 vs 24 hours wakefulness) affects attention (DV: score on a reaction-time test).
- A classroom experiment where two teaching methods are compared using random assignment of students to groups.
- A within-subjects study where students take both types of revision methods in different weeks and their recall is tested.
- Independent variable (IV) = variable manipulated.
- Dependent variable (DV) = measured outcome.
- Control group = group not receiving experimental treatment.
- Experimental group = group receiving the manipulated variable.
Correlational Method
Purpose of correlational studies
Correlational research examines whether two or more variables change together. It measures the strength and direction of relationships without manipulating variables. Correlational designs are useful for studying naturally occurring variables, for initial exploration of possible relationships, and when experimental control is impractical or unethical.
Understanding correlation coefficients
The relationship between variables is summarised by a correlation coefficient, typically ranging from -1 to +1. Positive values indicate variables increase together; negative values indicate one increases as the other decreases. The absolute value indicates strength: values near 0 indicate weak or no linear relationship, while values near 1 or -1 indicate strong linear relationships.
Types and forms
Correlational research can be cross-sectional (measured at one time), longitudinal/prospective (measured across time to assess prediction), or partial (controlling statistically for third variables). Researchers use scatterplots to visualise relationships and correlation coefficients (such as Pearson’s r for interval data or Spearman’s rho for ordinal data) to summarise them numerically.
Limitations: correlation is not causation
A crucial limitation is that correlation does not imply causation. Three alternative explanations can produce a correlation: A causes B, B causes A (reverse causation), or a third variable C causes both A and B (confounding). Additionally, correlations detect linear relationships; two variables may have a strong non-linear relationship with a low linear correlation coefficient.
Use in psychology
Correlational studies are widely used in personality psychology, epidemiology, educational research and health psychology to identify risk factors, predictors and associations that may guide further experimental or intervention work. For example, correlation can highlight that higher stress is associated with poor sleep, suggesting avenues for intervention research.
Reporting and interpreting
Reports should state the correlation coefficient, sample size, significance level and confidence intervals. Researchers should discuss effect sizes and practical importance and consider possible confounds. Where possible, longitudinal designs can provide stronger evidence about temporal order and potential causal directions, though they still cannot prove causation without further controls.
- Finding a positive correlation between study hours and exam scores in a sample of students.
- Observing a negative correlation between amount of exercise and reported stress levels.
- Using partial correlation to control for age when studying the link between social media use and sleep quality.
- Correlation coefficient (r): A statistical measure of linear association between two variables, ranging from -1 to +1.
- Positive correlation: r > 0. Negative correlation: r < 0.
- No linear correlation: r ≈ 0.
Survey Method and Questionnaires
What surveys measure
Surveys collect standardised information from samples to describe attitudes, beliefs, behaviours and background characteristics. They are efficient for gathering data from large groups and can be adapted to many formats: paper, online, telephone, or face-to-face interviews. Surveys allow comparisons across subgroups and can detect trends over time when repeated.
Designing questionnaires
Good questionnaire design is essential for valid results. Questions should be clear, concise and neutral. Avoid double-barrelled items (asking two things at once) and leading questions that push respondents toward a particular answer. Include instructions and examples where needed, and pilot-test the questionnaire to identify confusing items.
Types of questions and response formats
Questions can be closed-ended (multiple-choice, ranking, Likert scales) or open-ended. Closed items facilitate quantitative analysis; open items provide richer qualitative data but require coding. Likert scales (e.g., strongly agree to strongly disagree) are widely used for attitudes. For behaviour frequency, include clear time frames (e.g., 'in the past week'). Balanced response options with a neutral middle reduce forced choices, and using consistent scales across related items aids interpretation.
Sampling and administration
Surveys rely on representative sampling to generalise results. Random and stratified sampling improve representativeness; convenience samples limit generalisation. Mode of administration affects responses: anonymity may increase honesty on sensitive topics, while interviewer-administered surveys can help clarify questions but may introduce social desirability bias. Response rates matter—low response rates risk non-response bias, which should be considered and reported.
Data quality and piloting
Pilot testing helps identify ambiguous wording, estimate completion time and test the electronic or paper flow. Check for ceiling or floor effects where most respondents choose extreme options. Include attention checks for long questionnaires and use branching logic cautiously so respondents see only relevant items.
Ethics and sensitive topics
Ensure informed consent, confidentiality and the right to skip questions. For sensitive topics, anonymity and careful wording reduce harm and improve honesty. Offer information on support services if questions touch on distressing topics. Finally, report results honestly, noting limitations such as sampling bias or measurement error.
- Designing a short student survey on study habits using Likert items and a few open-ended questions for comments.
- Using an online questionnaire to measure attitudes toward career choices among higher secondary students.
- Piloting a survey on screen time to check whether respondents understand the time categories.
- Response bias: systematic error due to respondents answering inaccurately.
- Pilot testing: trial of a questionnaire on a small group to refine items.
Observation Methods
Overview of observation
Observation means systematically watching behaviour and recording it according to rules set before the study. It is especially useful when verbal report is impossible or unreliable, when studying natural interactions, or when the researcher wants to capture behaviour as it naturally occurs. Observation can provide rich detail and context about social behaviours, interactions, and routines.
Types of observation
Observational methods vary by setting and awareness. Naturalistic observation occurs in everyday settings without interfering, giving high ecological validity. Structured observation creates specific situations where behaviour can be compared across participants. Observers may be overt (participants know they are being observed) or covert (they do not). Each choice involves trade-offs between ethical transparency and reactivity (changing behaviour because of being observed).
Recording techniques
Accurate recording requires clear operational definitions of target behaviours. Common techniques include checklists (tick boxes for observed behaviours), rating scales (judging intensity or frequency), time sampling (recording behaviour at fixed intervals), and event sampling (noting every occurrence of a defined event). Video or audio recording enables later review and more detailed coding, though it raises privacy concerns and may require special consent.
Reliability and observer training
Observer bias and inconsistency are risks. Training observers with clear coding manuals, practice sessions, and calibration exercises improves reliability. Inter-observer reliability is assessed by having multiple observers code the same events independently and comparing their records. High agreement increases confidence that observations reflect the behaviour rather than individual observers’ judgments.
Reducing reactivity and bias
Reactivity can be reduced by prolonged presence (so participants become used to being observed), using unobtrusive measures, or covert observation when ethically permissible. Blinding observers to study hypotheses can reduce expectancy effects. Using objective, narrowly defined behavioural categories (e.g., 'number of times child shares toy') reduces subjective judgement.
Applications and limitations
Observation is used in child development, educational research, clinical assessment and organisational studies. It excels at describing complex social interactions and contextual factors but is often time-consuming and may lack generalisability if based on small samples. Ethical issues, especially regarding consent and privacy, must be handled carefully. Observational data are best combined with other methods for a fuller understanding of phenomena.
- Naturalistic observation of playground interactions to record instances of sharing among children.
- Structured observation in a lab where participants are placed in a problem-solving task and cooperative behaviours are recorded.
- Using time sampling to record how often a student raises their hand during a lesson at 5-minute intervals.
- Time sampling: recording behaviours at predetermined intervals.
- Event sampling: recording every occurrence of a targeted behaviour.
Case Study Method
Definition and purpose
A case study is an intensive, detailed investigation of a single individual, group, institution or event. It aims to gather rich qualitative and sometimes quantitative data to understand complex phenomena, unusual conditions, or developmental processes that are difficult to capture through large-scale studies.
Sources of data
Case studies use multiple information sources to build a comprehensive picture: clinical interviews, standardised tests, direct observations, archival records, diaries, and reports from family or teachers. Using varied sources increases the depth and credibility of findings, allowing the researcher to triangulate information and see patterns across different types of evidence.
Types and uses
Clinical case studies explore psychological disorders and therapy outcomes, while biographical case studies examine life histories and turning points. Single-case experimental designs apply experimental logic to one case, using repeated measurement and manipulation to assess cause-effect relations within an individual. Case studies are valuable for developing hypotheses, illustrating points in teaching, and informing practice in clinical or educational settings.
Strengths
They provide detailed, context-rich data that can reveal processes, mechanisms and rare phenomena that large surveys may miss. Case studies can suggest new theories, identify variables worth testing experimentally, and provide real-world insights for practitioners.
Limitations and generalisability
Because they focus on one case, generalising to wider populations is limited. Findings might reflect unique personal history, situational factors or researcher interpretation. Case studies can be influenced by researcher bias, selective reporting, and retrospective distortion in life histories. Researchers should be cautious in drawing broad claims from single cases.
Ethics and confidentiality
Because cases can be identifiable, protecting privacy is crucial. Obtain informed consent, anonymise details, and handle sensitive information with care. When publishing, change names and context details while preserving necessary information for understanding the case.
Reporting a case study
Good reports clearly outline methods, data sources, timeline, interpretations and limitations. They should separate observed facts from speculative interpretation and suggest how the case informs future research or practice. Case studies often conclude with practical implications and recommendations for further systematic study.
- A clinical case study following a patient with a rare cognitive impairment over several years, documenting tests and therapy responses.
- An educational case study examining a single school that implemented a new teaching method and tracking student performance and teacher reports.
Longitudinal and Cross-sectional Designs
Overview of designs
Longitudinal and cross-sectional designs are two common approaches for studying change over time. Both aim to describe developmental patterns or age differences, but they differ in whether the same individuals are followed or new samples are compared at one point in time.
Longitudinal design explained
In longitudinal studies the same participants are measured repeatedly across months, years or decades. This design can reveal how individuals change, identify patterns of stability and growth, and examine how earlier measures predict later outcomes. Because the same people are assessed over time, longitudinal research can suggest temporal sequences and potential causal directions more convincingly than a single snapshot.
Advantages of longitudinal research
It allows direct observation of developmental trajectories, helps separate age-related change from individual differences, and supports prediction of long-term outcomes from early measures. It is useful in developmental psychology, education and health research where long-term effects matter.
Challenges of longitudinal research
Long studies require substantial resources and sustained participant cooperation. Attrition—the loss of participants over time—can bias results if those who drop out differ systematically from those who remain. Repeated testing can produce practice effects or sensitisation where participants' responses change because they become familiar with measures. Researchers must plan for these issues through careful design and statistical methods to handle missing data.
Cross-sectional design explained
Cross-sectional research measures different age or cohort groups at a single time and compares them. It is quicker and less expensive than longitudinal work and avoids problems of attrition or repeated testing. Cross-sectional studies are useful for initial comparisons of groups, to describe age-related differences, or to test hypotheses about group contrasts.
Limitations of cross-sectional studies
Cross-sectional differences may reflect cohort effects—differences in life experiences, education or culture—rather than developmental change. For example, older participants may differ from younger ones because of generational factors, not aging itself. Thus cross-sectional findings require cautious interpretation.
Combining approaches and sequential designs
Mixed approaches, such as sequential designs, combine elements of both to disentangle age, period and cohort effects. Where possible, researchers choose the design that best fits the question: longitudinal designs for developmental processes, cross-sectional for quick comparisons. Reporting limitations and considering complementary methods strengthens conclusions.
- A longitudinal study that follows a group of children from primary school through adolescence to study the development of self-esteem.
- A cross-sectional study comparing memory performance in 10-, 20- and 60-year-olds to assess age-related differences.
Psychological Testing and Measurements
What are psychological tests?
Psychological tests are standardised tools designed to measure abilities, traits, symptoms or attitudes. They translate psychological constructs into measurable scores that can be compared across people. Tests can be used in education, clinical assessment, personnel selection and research.
Standardisation and norms
Standardisation means the test is administered and scored consistently for all individuals. Test manuals describe standard procedures and produce normative data—scores from a representative sample that allow interpretation of an individual’s results (e.g., percentile ranks, standard scores). Norms help determine whether a score is typical, high or low relative to a reference group.
Types of tests
Common categories include ability tests (measuring intelligence or specific cognitive skills), achievement tests (measuring learned knowledge), personality inventories (assessing traits and preferences), and clinical instruments (screening for mental health symptoms). Tests may be objective (fixed-choice items) or more subjective and projective in nature.
Reliability and validity in tests
Reliability refers to the consistency of test scores: a reliable test yields similar results under consistent conditions. Forms include test-retest stability, inter-rater agreement, and internal consistency across items. Validity concerns whether the test measures what it claims. Content validity checks coverage of the domain; criterion validity examines relation to external criteria; construct validity examines theoretical relationships. A test must be reliable before it can be valid, but reliability alone does not ensure validity.
Practical considerations and fairness
Selecting a test involves checking its appropriateness for age, language and cultural background. Tests must be used responsibly: administrators need training in giving tests, scoring and interpreting results. Cultural bias and language differences can affect scores; test users should consider fairness and whether local norms are available.
Ethical use and feedback
Testing carries ethical responsibilities: obtain informed consent, maintain confidentiality, provide feedback that is understandable and useful, and avoid misusing scores (e.g., making major life decisions based on a single test without corroborating evidence). Tests should be one part of a broader assessment process.
- Interpreting a standard score by comparing a student’s result to the normative sample provided in the test manual.
- Checking the test-retest reliability of a short anxiety scale by administering it twice two weeks apart and calculating the correlation.
- Reliability coefficient (r): measure of consistency, ranges from 0 to 1.
- Validity: degree to which test measures intended construct; no single formula, assessed by multiple methods.
Sampling Techniques
Why sampling matters
Researchers rarely study entire populations. Instead, they select samples—subsets expected to represent the larger group. Good sampling enables generalisation of results and helps researchers make inferences about populations with known uncertainty. Poor sampling produces biased results that may mislead policy and practice.
Probability sampling methods
Probability sampling gives every member of the population a known chance of selection. Simple random sampling uses random numbers to select participants, ensuring equal chance and strong statistical foundations. Stratified sampling divides the population into meaningful subgroups (strata) such as age, gender or school grade, and then randomly samples within each stratum to preserve subgroup proportions. Systematic sampling selects every nth person from a list after a random start; it is easy to implement but depends on a reliable sampling frame.
Non-probability sampling methods
Non-probability sampling does not give known selection probabilities and is often easier and cheaper. Convenience sampling recruits those who are easily available, while quota sampling fills set quotas of subgroups but without random selection. Purposive sampling targets people with specific characteristics, useful in qualitative or exploratory research. Snowball sampling asks participants to refer others and is helpful for hard-to-reach populations. While practical, these methods limit generalisability and require careful acknowledgement of bias.
Sampling frame and errors
The sampling frame is the list from which the sample is drawn. Errors can arise if the frame excludes segments of the population or is outdated. Non-response bias occurs when those who decline to participate differ systematically from those who respond. Researchers should report response rates and consider weighting or follow-up methods to address bias.
Sample size considerations
Sample size affects precision: larger samples generally provide more reliable estimates and greater power to detect effects. Required size depends on expected effect size, acceptable error margin and available resources. Pilot studies can inform realistic sample size estimates and recruitment strategies.
Choosing a method
Choice depends on the study goal: for descriptive or inferential studies aiming for population estimates, probability sampling is preferred. For exploratory, qualitative or constrained studies, non-probability samples may be acceptable with appropriate caution. Clear reporting of sampling decisions and limitations helps readers interpret the findings accurately.
- Using stratified sampling to ensure a school study includes students from each grade level in proportion to enrolment numbers.
- Selecting participants by convenience for a pilot classroom study, while noting limitations in generalising the results.
Measurement Scales and Levels of Data
Introduction to measurement levels
Data collected in psychological research come in different levels of measurement. These levels—nominal, ordinal, interval and ratio—determine which statistical methods are appropriate and how results should be interpreted. Choosing the right summary statistics and tests depends on recognising a variable’s level.
Nominal scale
Nominal data are categories without order, such as gender, blood type or favourite colour. Analysis involves counts and proportions; the mode is a meaningful central measure. Nominal data cannot be averaged or ranked in a meaningful way.
Ordinal scale
Ordinal data indicate order but not equal intervals. Examples include rank order, educational levels, or Likert-type ratings (agree to disagree). With ordinal data, median and percentiles are appropriate summaries. Differences between ranks are not necessarily equal, so calculating means may be misleading without caution.
Interval scale
Interval data have equal intervals between values but lack a true zero point. A classic example is temperature in Celsius. Differences are meaningful but ratios are not (e.g., 20°C is not twice as hot as 10°C). Interval data allow use of means, standard deviations, and many parametric tests if other distributional assumptions hold.
Ratio scale
Ratio data have equal intervals and a meaningful zero, like reaction times, counts of correct answers, or weight. All arithmetic operations, including ratios, are valid. Ratio data permit the full range of statistical analyses and meaningful statements about how many times larger one value is compared to another.
Implications for analysis
Parametric tests usually assume interval or ratio data and often normal distributions; non-parametric tests are used for nominal or ordinal data or when assumptions are violated. Understanding data levels guides choice of summary statistics, graphs, and statistical tests. For multi-item scales, researchers check whether items produce interval-like scores before applying parametric methods.
Transformations and composite scores
Researchers sometimes create composite scores by summing items; they should ensure items measure the same construct and check internal consistency. Data transformation can improve normality for parametric tests, but transformations change interpretability and must be justified. Clear documentation of measurement levels and transformations helps readers assess the appropriateness of analyses.
- Classifying responses to a multiple-choice question as nominal data.
- Using median rather than mean to summarise ordinal satisfaction ratings.
- Levels of measurement: Nominal < Ordinal < Interval < Ratio
- Mean, median, mode depend on data level and distribution.
Reliability and Validity
Core concepts defined
Reliability and validity are central to assessing the quality of measures and research. Reliability refers to the consistency and stability of measurement: a reliable instrument yields similar results under consistent conditions. Validity refers to whether an instrument measures what it intends to measure. A measure must be reliable to be valid, but a reliable measure is not necessarily valid.
Types of reliability
Common forms of reliability are:
- Test-retest reliability: stability of scores over time when the construct should be stable.
- Inter-rater reliability: agreement between different observers or raters coding the same events.
- Internal consistency: how well items in a scale measure the same construct, often estimated by Cronbach’s alpha.
High reliability reduces measurement error and increases the likelihood that observed differences reflect true differences rather than random noise.
Types of validity
Validity has several facets:
- Content validity: the extent to which a test covers the full range of the concept (e.g., a depression scale covering mood, sleep, appetite).
- Criterion validity: how well test scores relate to an external criterion, either concurrently or predictively (e.g., an aptitude test predicting job performance).
- Construct validity: whether the test behaves as theory predicts, including convergent validity (correlates with related measures) and discriminant validity (does not correlate with unrelated measures).
Assessing and improving reliability and validity
Researchers estimate reliability coefficients (e.g., test-retest correlations, Cronbach’s alpha) and present validity evidence through correlations, factor analysis or criterion studies. Improving reliability involves clearer items, standardised administration and training raters. Strengthening validity includes careful operationalisation, using multiple methods (triangulation), and ensuring items reflect the construct comprehensively.
Trade-offs and reporting
High reliability does not guarantee validity: a consistently biased measure can be reliable but invalid. Researchers should report reliability and validity evidence in study methods so readers can judge the quality of measures. Transparent reporting, including limitations, supports accurate interpretation of findings.
- Estimating test-retest reliability for a mood scale by correlating scores from two administrations two weeks apart.
- Demonstrating construct validity by showing a new anxiety scale correlates highly with an established anxiety measure (convergent validity) and not with an unrelated trait (discriminant validity).
- Internal consistency: Cronbach’s alpha (α) used to estimate how closely related a set of items are as a group.
- Reliability coefficient: r, ranges from 0 (no reliability) to 1 (perfect reliability).
Ethics in Psychological Research
Why ethics matter
Ethical standards protect participants, uphold scientific integrity and maintain public trust. In psychology, ethical rules guide how people are treated during research, balancing the pursuit of knowledge with respect for individual rights and welfare. Ethical research minimises harm, ensures voluntary participation, and treats data responsibly.
Core ethical principles
Key principles include informed consent (participants decide to join with full understanding), confidentiality (personal data protected), beneficence (maximising benefits), non-maleficence (avoiding harm), and respect for persons (recognising autonomy and vulnerability). These principles guide everyday research decisions, from recruitment language to data storage.
Informed consent and assent
Informed consent requires clear information about the study’s purpose, procedures, risks and benefits, and the right to withdraw without penalty. For minors or those unable to consent, parental or guardian consent is required along with the participant’s assent if age-appropriate. Consent is an ongoing process, and participants should be informed of new information that could affect willingness to continue.
Deception and debriefing
Deception—providing misleading or incomplete information—may sometimes be used if it is necessary and no alternatives exist. Ethical guidelines require that deception be minimised, justified by potential benefits, and followed by a thorough debriefing that explains the true purpose, corrects misconceptions, and offers support if needed. Ethics committees evaluate whether deception is acceptable in each case.
Privacy, confidentiality and data protection
Researchers must store data securely, limit access, and report results in ways that prevent identification of individuals. Anonymity is ideal for sensitive surveys; when identifiable data are needed, strong safeguards and clear participant information are required. Local laws and institutional rules on data protection must be followed.
Special populations and additional safeguards
Research involving children, patients, prisoners or those with impaired decision-making requires extra protections, such as independent consent procedures, assessment of capacity, and monitoring of welfare. Vulnerable participants should not be exposed to undue risk for the sake of research benefits to others.
Ethical review and responsibility
Most research requires review by an ethics committee or institutional review board, which assesses risks, consent procedures and safeguards. Researchers must report adverse events and adhere to approved protocols. Ethical conduct also means honest reporting, avoiding fabrication or selective reporting, and crediting contributors appropriately.
- Obtaining parental consent and child assent before a classroom study on learning strategies.
- Debriefing participants after a memory experiment that involved misleading information about the study’s purpose.
Research Design: Variables and Controls
Understanding variables
Variables are the elements that vary in research and form the basis of hypotheses and testing. Independent variables (IVs) are manipulated or defined as predictors; dependent variables (DVs) are the outcomes measured. Other important types include control variables (kept constant), extraneous variables (which may affect DVs), and confounding variables (extraneous factors that are correlated with the IV and may distort conclusions).
Operationalisation and clarity
Operationalisation translates abstract constructs into measurable or manipulable forms. Clear operational definitions reduce ambiguity and support replication. For example, defining 'attention' as score on a continuous performance task gives a concrete DV that can be measured reliably.
Control strategies
Control reduces alternative explanations. Common techniques include randomisation (to evenly distribute unknown factors across groups), matching (pairing participants on key traits), holding conditions constant (same environment and instructions), and counterbalancing (to address order effects in within-subjects designs). Studies often include control groups or placebo conditions to separate specific effects of an intervention from general effects like practice or expectation.
Quasi-experimental designs
When random assignment is not feasible or ethical, quasi-experiments use naturally occurring groups or conditions. While they resemble true experiments, quasi-experiments require careful design and statistical control to address selection biases and confounds. Researchers may use matching, statistical controls or interrupted time series designs to strengthen causal inferences.
Manipulation checks and pilot studies
Manipulation checks verify that the IV was perceived or experienced as intended (for example, confirming that a mood induction changed participants' mood). Pilot studies test procedures, instruments and timing on a small scale, revealing logistical problems and improving measures before the main study.
Documenting design choices
Transparent reporting of variables, operational definitions, control methods and limitations allows readers to assess internal validity and potential biases. Discussing alternative explanations and how the design addressed them strengthens the study’s credibility.
- Using counterbalancing to control order effects in a within-subjects design testing two study techniques.
- Including a placebo group to control for expectancy effects in a study of a new relaxation intervention.
Data Collection and Recording
Collecting data systematically
Data collection is the careful gathering of information using tools and procedures decided during study planning. Methods include tests, questionnaires, structured interviews, observations, physiological measures and digital logs. A clear protocol describes who collects data, when and how, ensuring consistency and reducing errors.
Recording formats and codebooks
Data should be recorded in organised formats with clear labels and codebooks that define variable names, units and permitted values. Codebooks help anyone who later works with the dataset to understand the meaning of each column and the coding scheme for categorical variables. Electronic data entry should include validation rules to prevent impossible values (e.g., age = -5).
Quality control measures
Quality control reduces errors during collection and entry. Techniques include training data collectors, double data entry, random audits, inter-rater reliability checks for coded data, and automated validation checks in electronic forms. Video or audio records can serve as backups for observation studies and facilitate reliability checks.
Dealing with missing data
Missing data arise for many reasons: skipped items, dropout, or equipment failure. Researchers should plan to minimise missingness through clear instructions, reminders and good equipment. When missing data occur, options include analysing only complete cases, using statistical imputation methods to estimate missing values, or using models that accommodate missingness. The chosen approach must be justified and sensitivity analyses reported to show how results depend on missing data handling.
Data security and confidentiality
Protecting participant information is essential. Store identifiable data securely with restricted access, use encryption where appropriate, and separate identifying information from research data. Anonymise or pseudonymise datasets before sharing, and follow institutional and legal requirements for data retention and disposal.
Preparing for analysis
Before analysis, clean the data by checking for outliers, inconsistent responses, and coding errors. Create derived variables with clear documentation. Maintain a reproducible workflow with scripts that show how raw data were transformed into analysis datasets. Good documentation supports transparency and replication.
- Creating a coding sheet for categorising responses to an open-ended question and training two coders to apply it.
- Using a spreadsheet with validation rules to prevent impossible numerical entries (e.g. negative ages).
Basic Data Analysis and Interpretation
Descriptive statistics
Descriptive statistics summarise and describe data. Measures of central tendency (mean, median, mode) indicate typical scores, while measures of dispersion (range, variance, standard deviation) describe spread. Frequency tables, histograms and bar charts visualise distributions and highlight skew, multimodality or outliers. Clear labelling of axes and units is essential for interpretable graphs.
Inferential statistics basics
Inferential statistics help decide whether observed patterns in a sample likely reflect real effects in the population. Common tests introduced at this level include t-tests for comparing two group means, chi-square tests for categorical data, and correlation coefficients for association between variables. Tests yield p-values indicating the probability of observing data as extreme as those collected if the null hypothesis were true. Researchers also report effect sizes and confidence intervals to communicate the magnitude and precision of effects.
Choosing appropriate tests
Choice of test depends on measurement level (nominal, ordinal, interval, ratio), sample size and assumptions such as normality and equal variances. Parametric tests assume interval/ratio data and particular distributions; non-parametric tests are alternatives for ordinal data or when assumptions are violated. Checking assumptions through plots and tests informs the correct analytic approach.
Interpreting results responsibly
Statistical significance does not equal practical importance. A small effect can be statistically significant in a large sample but may have limited real-world impact. Conversely, a large effect may not reach significance in a small sample. Reporting both statistical results (test statistics, p-values) and effect sizes helps readers assess relevance. Discussing limitations such as sampling bias or measurement error is essential.
Avoiding common errors
Common mistakes include multiple testing without adjustment (increasing false positives), p-hacking (selective reporting to find significant results), and overgeneralising from non-representative samples. Transparency about analysis choices, pre-registration where possible, and reporting all relevant findings reduce these risks.
Presenting results
Present results with clear tables and figures that summarise key statistics. Use visuals to complement but not replace concise textual explanations. Conclusions should align with the evidence, acknowledge uncertainty, and suggest directions for further study or practical implications.
- Calculating the mean and standard deviation for test scores in two groups and plotting them in a bar chart with error bars.
- Computing a correlation coefficient between study hours and marks and interpreting its magnitude and direction.
- Mean = sum of scores / number of scores.
- Standard deviation (SD): square root of the variance.
- Correlation coefficient r quantifies linear association between two variables.
Reporting Research and Critical Evaluation
Purpose of reporting
Research reporting shares methods, data and conclusions so others can evaluate, replicate and build on findings. Clear reporting advances science by allowing independent scrutiny and by providing a record of what was done and why. Reports vary in audience and format but should maintain transparency about methods and limitations.
Structure of research reports
Standard sections include an introduction (context and hypotheses), method (participants, materials, procedure, analysis), results (descriptive and inferential findings), and discussion (interpretation, limitations, implications). Appendices can include instruments, coding schemes and additional analyses. The method section should be detailed enough to allow replication.
Presenting results clearly
Results should include both descriptive statistics (means, SDs, frequencies) and inferential results (test statistics, p-values, effect sizes, confidence intervals). Tables and figures should be self-explanatory with clear titles, labels, units and legends. Avoid overloading visuals; each table or graph should communicate a focused message.
Critical evaluation of research
Evaluating a study involves assessing sampling (representativeness, size), measurement quality (reliability, validity), design (controls, randomisation), analysis (appropriate tests, handling of missing data), and ethics. Consider whether conclusions are supported by data and whether alternative explanations were considered. Good evaluations note strengths and limitations and suggest improvements or further tests.
Writing abstracts and summaries
An abstract provides a concise summary of background, methods, key results and main conclusions. It helps readers decide whether to read the full report. For wider audiences, write a brief plain-language summary that highlights practical implications and caveats without technical jargon.
Peer review and publication
Peer review is a quality-check process where other experts evaluate methods, analyses and claims before publication. Reviewers check for methodological rigor, appropriate analysis and ethical compliance. Peer-reviewed publication increases credibility, but readers should still appraise each study critically for design and context.
Reproducibility and transparency
Transparent reporting includes sharing methods, materials and anonymised data when possible. Pre-registration of hypotheses and analysis plans reduces selective reporting. Reproducibility—other researchers obtaining similar results using the same methods—strengthens confidence in findings. Where replication fails, consider differences in samples, settings, or measurement that might explain divergence.
Referencing and avoiding plagiarism
Always credit prior work through proper citation. Use a consistent referencing style and include complete details so others can locate sources. Avoid copying text or ideas without attribution; accurate referencing demonstrates scholarly honesty and situates new findings within existing knowledge.
Communicating limitations and implications
A responsible discussion balances enthusiasm about findings with clear statements of limitations: sample size, sampling method, measurement weaknesses, possible confounds and generalisability. Suggest realistic implications and next research steps rather than overstating claims. For policy or practice recommendations, explain the strength of evidence and any conditions under which findings apply.
Dissemination beyond academia
Researchers may present findings to schools, policy-makers or the public. Tailor messages to the audience: use infographics, short summaries or presentations with clear recommendations. When communicating outside academia, avoid sensationalist language and explain uncertainty plainly.
- Writing a short report of a school experiment including aim, method, results with a bar chart, and conclusion.
- Critically evaluating a published study by listing its strengths and limitations in sampling and measurement.
Key Concepts
- Hypothesis
- A clear, testable prediction about the relationship between variables.
- Independent Variable
- The variable that is manipulated or varied by the researcher.
- Dependent Variable
- The outcome measured to see the effect of the independent variable.
- Operationalisation
- Specifying exactly how a variable will be measured or manipulated.
- Random Sampling
- A sampling method where every member of the population has an equal chance of selection.
- Random Assignment
- Allocating participants to groups by chance to reduce bias.
- Reliability
- The consistency of a measurement across time, observers, or items.
- Validity
- The extent to which a test measures what it claims to measure.
- Correlation
- A statistical measure of the relationship between two variables.
- Case Study
- An in-depth study of a single person, group or event using multiple data sources.
- Ethics
- Standards to protect participants from harm and to ensure research integrity.
- Standardisation
- Administering procedures and measures in a consistent way for all participants.
- Sampling Bias
- Systematic error introduced when a sample is not representative of the population.
- Confounding Variable
- An outside factor that may influence the dependent variable and distort results.
- Pilot Study
- A small trial study conducted to test procedures and instruments before the main study.
- Descriptive Statistics
- Numbers that summarise data, such as mean, median and standard deviation.
- Inferential Statistics
- Statistical tests used to draw conclusions about a population from sample data.
Practice Questions
-
What is the purpose of operationalising variables in a study? / किसी अध्ययन में चर को ऑपरेशनलाइज़ करने का क्या प्रयोजन है?
Show answer
Operationalising variables makes abstract concepts measurable and specifies how they will be observed or manipulated, enabling clear measurement and replication. / चर को ऑपरेशनलाइज़ करने का उद्देश्य अमूर्त अवधारणाओं को मापनीय बनाना और यह स्पष्ट करना है कि उन्हें कैसे मापा या नियंत्रित किया जाएगा, जिससे मापन स्पष्ट और अध्ययन की नकल संभव हो जाती है।
-
Explain the difference between an experiment and a correlational study. / एक प्रयोग और सहसंबंध अध्ययन में क्या अंतर है समझाइए।
Show answer
An experiment manipulates an independent variable and uses control to test causal effects on a dependent variable, often with random assignment. A correlational study measures two or more variables to see if they are related but does not manipulate variables and cannot establish causation. / एक प्रयोग में स्वतंत्र चर को बदला जाता है और नियंत्रित परिस्थितियों में परिणामस्वरूप होने वाले प्रभावों का परीक्षण किया जाता है, अक्सर यादृच्छिक आवंटन के साथ; जबकि एक सहसंबंध अध्ययन दो या अधिक चरों को मापता है यह देखने के लिए कि वे संबंधित हैं या नहीं, पर इसमें चरों का हेरफेर नहीं होता और यह कारण-संबंध स्थापित नहीं कर सकता।
-
List three ways to reduce observer bias during observation. / अवलोकन के दौरान पर्यवेक्षक पूर्वाग्रह को कम करने के तीन तरीके बताइए।
Show answer
1) Train observers and use clear operational definitions. 2) Use multiple observers and calculate inter-observer reliability. 3) Use blind or covert observation or video recording for later coding. / 1) पर्यवेक्षकों को प्रशिक्षित करें और स्पष्ट ऑपरेशनल परिभाषाएँ उपयोग करें। 2) कई पर्यवेक्षकों का उपयोग करें और अंतर-पर्यवेक्षक विश्वसनीयता निकालें। 3) अंधा या गुप्त अवलोकन करें या बाद में कोडिंग के लिए वीडियो रिकॉर्डिंग का उपयोग करें।
-
What is the null hypothesis and why is it important in statistical testing? / शून्य परिकल्पना क्या है और सांख्यिकीय परीक्षण में यह क्यों महत्वपूर्ण है?
Show answer
The null hypothesis states that there is no effect or difference. It provides a default position that statistical tests evaluate; rejecting the null suggests evidence for the alternative hypothesis. It is important because it offers a clear criterion for decision-making based on probability. / शून्य परिकल्पना यह कहती है कि कोई प्रभाव या अंतर नहीं है। यह एक डिफ़ॉल्ट स्थिति प्रदान करती है जिसे सांख्यिकीय परीक्षण मूल्यांकन करते हैं; शून्य को ठुकराने का अर्थ विकल्प परिकल्पना के लिए साक्ष्य मिलना है। यह निर्णय-निर्धारण के लिए संभाव्यता के आधार पर स्पष्ट मानदंड देती है।
-
Describe stratified sampling and give one advantage. / स्तरीकृत नमूनाकरण का वर्णन कीजिए और इसका एक लाभ बताइए।
Show answer
Stratified sampling divides the population into subgroups (strata) based on characteristics, then randomly samples from each stratum in proportion to its size. An advantage is increased representativeness of key subgroups, reducing sampling error. / स्तरीकृत नमूनाकरण में जनसंख्या को विशेषताओं के आधार पर उपसमूहों में विभाजित किया जाता है और फिर प्रत्येक उपसमूह से उसके आकार के अनुपात में यादृच्छिक रूप से नमूना लिया जाता है। इसका लाभ यह है कि महत्वपूर्ण उपसमूहों का प्रतिनिधित्व बेहतर होता है, जिससे नमूना त्रुटि घटती है।
-
A researcher finds r = 0.65 between hours of practice and performance. How should this be interpreted? / एक शोधकर्ता अभ्यास घंटों और प्रदर्शन के बीच r = 0.65 पाता है। इसे कैसे व्याख्यायित किया जाना चाहिए?
Show answer
An r of 0.65 indicates a moderately strong positive linear relationship: as practice hours increase, performance tends to increase. It does not prove causation and other factors might contribute. Strength also depends on sample size and context. / r = 0.65 एक मध्यम से मजबूत सकारात्मक रैखिक संबंध दर्शाता है: जैसे-जैसे अभ्यास के घंटे बढ़ते हैं, प्रदर्शन भी बढ़ने की प्रवृत्ति होती है। यह कारण-संबंध साबित नहीं करता और अन्य कारक भी योगदान दे सकते हैं। ताकत का अर्थ नमूना आकार और संदर्भ पर भी निर्भर करता है।
-
Why is informed consent necessary, and what special steps are taken when participants are minors? / सूचित सहमति आवश्यक क्यों है, और जब प्रतिभागी नाबालिग हों तो क्या विशेष कदम उठाए जाते हैं?
Show answer
Informed consent ensures participants know the study purpose, procedures, risks, benefits and their rights, allowing voluntary participation. With minors, parental or guardian consent is required along with the child’s assent; procedures must be age-appropriate and extra care taken to protect welfare. / सूचित सहमति सुनिश्चित करती है कि प्रतिभागी अध्ययन का उद्देश्य, प्रक्रियाएँ, जोखिम, लाभ और उनके अधिकार जानते हैं, जिससे स्वैच्छिक भागीदारी संभव होती है। नाबालिगों के साथ माता-पिता या अभिभावक की सहमति आवश्यक होती है साथ ही बच्चे की सहमति (assent) भी ली जाती है; प्रक्रियाएँ आयु के अनुकूल होनी चाहिए और कल्याण की अतिरिक्त सुरक्षा आवश्यक है।
-
Explain the difference between reliability and validity with an example. / विश्वसनीयता और वैधता के बीच अंतर एक उदाहरण के साथ समझाइए।
Show answer
Reliability is consistency; validity is accuracy. For example, if a stopwatch gives the same time each trial for a task it is reliable. If the stopwatch is stopped late every time, it is reliable but not valid for measuring true task duration. A valid measure would record the correct duration consistently. / विश्वसनीयता निरंतरता है; वैधता सटीकता है। उदाहरण के लिए, यदि एक स्टॉपवॉच किसी कार्य के हर प्रयास के लिए एक समान समय देता है तो वह विश्वसनीय है। अगर वह हर बार देर से रोका जाता है तो वह विश्वसनीय तो होगा पर सच्ची कार्य अवधि मापने के लिए वैध नहीं है। एक वैध मापक वह होगा जो सही अवधि को लगातार रिकॉर्ड करे।
-
Give two reasons why pilot testing is useful before the main study. / मुख्य अध्ययन से पहले पायलट परीक्षण उपयोगी होने के दो कारण बताइए।
Show answer
Pilot testing helps identify unclear questions, procedural problems or technical issues and estimates time and response rates, allowing improvements before full data collection. It also checks whether measures behave as expected. / पायलट परीक्षण अस्पष्ट प्रश्नों, प्रक्रियात्मक समस्याओं या तकनीकी मुद्दों की पहचान करने और समय तथा प्रतिक्रिया दरों का अनुमान लगाने में मदद करता है, जिससे पूर्ण डेटा संग्रह से पहले सुधार किए जा सकते हैं। यह यह भी जाँचता है कि माप अपेक्षित रूप से कार्य कर रहे हैं या नहीं।
-
What is a manipulation check and when would you use it? / मैनिपुलेशन चेक क्या है और आप कब इसका उपयोग करेंगे?
Show answer
A manipulation check is a test to confirm that the independent variable was experienced or perceived by participants as intended (for example, that a mood induction actually changed mood). It is used when researchers need to verify that their manipulation worked before interpreting effects on the dependent variable. / मैनिपुलेशन चेक एक परीक्षण है जो यह पुष्टि करता है कि स्वतंत्र चर प्रतिभागियों द्वारा इच्छानुसार अनुभव या प्रतिपादित किया गया था (उदाहरण के लिए, कि मनोदशा प्रेरणा ने वास्तव में मनोदशा बदली)। इसका उपयोग तब किया जाता है जब शोधकर्ता यह सत्यापित करना चाहते हैं कि प्रभावशीलता को व्याख्यायित करने से पहले उनकी मैनिपुलेशन सफल रही।
Related Laws & Principles
Explore allFoundational laws & principles behind this chapter. Each one opens a full page — what it says, why it matters, five practice questions and the mistakes to avoid.