No history yet

Evaluate Study Design

Trust and Applicability

Every research paper makes a claim. Your job is to decide if that claim is believable. This boils down to two fundamental questions: Can you trust the results? And do they apply to the real world? In research, these concepts are called internal and external validity.

[{}] asks if the study was done right. It's about whether the conclusions are warranted by the methods. Did the intervention truly cause the outcome, or was something else responsible?

External validity, or generalizability, asks if the results matter outside the study's specific context. Can you apply these findings to a broader population or setting?

A study can have high internal validity but low external validity. For example, a drug trial conducted on a very specific group of non-smoking, 20-25 year old male athletes might produce very clean, reliable results for that group. But those findings might not apply to an elderly woman with multiple health conditions. Achieving a balance between the two is one of the toughest challenges in research design.

Looking Back vs. Moving Forward

The timing of data collection is a critical design choice. A study can either look backward in time (retrospective) or follow participants forward (prospective).

Retrospective studies, like case-control studies, start with an outcome and look back for exposures. They are efficient for studying rare diseases and are relatively quick and inexpensive. However, they are vulnerable to biases. The most significant is —people with a condition may remember past exposures differently than those without it. Data quality can also be a problem, as you're relying on records that weren't created for research purposes.

Prospective studies, such as cohort studies and , identify a group of people and follow them over time to see who develops an outcome. This approach is powerful. By collecting data in real-time, you reduce recall bias and can establish a clearer temporal relationship between exposure and outcome. The downside? They are expensive, take a long time, and aren't practical for diseases that take decades to develop.

The Power of Comparison

To know if a treatment works, you need something to compare it against. That's the role of the control group. The choice of control is a major decision that shapes what question the study can answer.

Control TypeDescriptionKey Question AnsweredMajor Weakness
PlaceboAn inert substance or sham procedure that looks identical to the active intervention.Does the treatment have a true biological effect beyond patient expectation?Can be unethical if an effective treatment already exists.
Active ControlThe new intervention is compared against a known, effective treatment (standard of care).Is the new treatment better than, or at least as good as, the current standard?Requires larger sample sizes to prove superiority or non-inferiority.
Historical ControlA group of patients treated in the past is used as the comparator.How do outcomes with this new treatment compare to past results?Patient populations and standards of care change over time, introducing significant bias.

A placebo control is the best way to determine a treatment's absolute efficacy. However, an active control often answers a more relevant clinical question: Should I use this new treatment instead of the one I'm already using? Historical controls are the weakest and are generally only used in preliminary studies or when it's unfeasible or unethical to have a concurrent control group.

The Evidence Hierarchy

Not all evidence is created equal. The hierarchy of evidence is a model for ranking different study designs based on how well they minimize bias. At the top are systematic reviews and meta-analyses, which synthesize the results of multiple high-quality studies. Below them are RCTs, followed by observational studies like cohort and case-control studies.

Lesson image

This pyramid isn't a rigid rulebook. A well-conducted cohort study can sometimes provide more useful evidence than a poorly designed or implemented RCT. The hierarchy is a guide to help you assess the likely internal validity of a study design. When you see a case report or expert opinion, you should be far more skeptical than when you see a meta-analysis of several large RCTs.

The final piece of the puzzle is the study population itself, defined by its inclusion and exclusion criteria. Very narrow criteria—for example, only including patients with a specific genetic marker—can increase a study's internal validity by creating a homogenous group. But it simultaneously limits the external validity, as the results might not apply to the wider patient population. Always check these criteria to understand exactly who the study results apply to.

Critical evaluation aids in identification of strengths and weaknesses of a study and its relevance and validity in the clinic1; relevant information for critical evaluation is typically found in the methods and results sections.

Ready to test your ability to appraise study design?

Quiz Questions 1/5

A clinical trial for a new hypertension drug only includes non-smoking male participants between the ages of 30 and 40 with no other health conditions. What is the most likely trade-off in this study's design?

Quiz Questions 2/5

What is the most significant type of bias that retrospective studies, such as case-control studies, are particularly vulnerable to?