On this page
The mistake almost never shows up when it's made. A vague research question gets discovered eight months later, at the analysis stage, when it turns out the "primary outcome" was never actually defined precisely enough to measure. A missing confounding variable gets discovered at manuscript writing, when a reviewer asks a question the dataset simply can't answer. By the time the mistake is visible, the free fix — five minutes of extra planning — is no longer available, and the only options left are expensive ones.
Almost every one of these mistakes falls into one of three stages: design, data, or writing. None of them require more intelligence to avoid — they require checking a short list of things before, not after, the point of no return.
Why This Happens
Research skills are assumed, not taught
Clinical training rarely includes a structured walkthrough of how to plan a study end-to-end. Most physicians learn research the way they learn a new procedure — by doing it, making mistakes, and hopefully having someone senior catch the important ones in time.
Deadlines compress the planning phase first
When time is short, the planning stage is usually what gets cut, since it produces no visible output. Ironically, it's the stage where a caught mistake costs the least.
Solo research means fewer checkpoints
A project run mostly independently, without a structured review at each stage, loses the natural check that a second person catches something the original researcher stopped noticing.
Common Mistakes We See
Design-stage mistakes
- A research question that isn't specific enough to measure. "Does drug X help patients with condition Y" isn't answerable until it specifies the population, comparison, and outcome (a PICO structure).
- No clearly defined primary outcome. Studies with several "co-primary" outcomes and no pre-specified priority make both analysis and reviewer evaluation harder.
- Sample size decided by convenience, not calculation. "However many patients we can realistically recruit" is a resourcing constraint, not a substitute for a power calculation.
Data-stage mistakes
- No data dictionary before collection starts. Deciding variable definitions and units as you go leads to inconsistent entries that are expensive to clean later.
- Key confounders not recorded. Discovered only once a reviewer asks whether the groups were adjusted for a variable nobody thought to collect.
Writing-stage mistakes
- Results and discussion sections blended together. The results section should report findings only; interpretation belongs in the discussion.
- Reporting guidelines ignored. CONSORT, STROBE, or PRISMA (depending on design) should shape the manuscript from the first draft, not get retrofitted before submission.
Practical Recommendations
| Problem | Fix |
|---|---|
| Vague research question | Write it in PICO format before drafting the proposal |
| No sample size justification | Run a formal power calculation before data collection begins |
| Inconsistent data entry | Build a data dictionary defining every variable before collection starts |
| Missing confounders | List variables a reviewer is likely to ask about, and confirm each is collected |
| Blended results/discussion | Draft results as findings only; move all interpretation to discussion |
| No reporting checklist used | Identify the correct checklist (CONSORT, STROBE, PRISMA) at the protocol stage |
Quick Checklist
- Research question is written in PICO format
- A single primary outcome is defined and pre-specified
- Sample size is calculated, not estimated by convenience
- A data dictionary exists before data collection starts
- Likely confounders are identified and confirmed as collected variables
- Results and discussion are kept as separate, distinct sections
- The correct reporting checklist (CONSORT, STROBE, PRISMA) is identified before writing begins
- A second reviewer has read the protocol before data collection starts
Need Expert Guidance?
Book a free call with PureMed and let our research specialists help you save valuable time, avoid costly research mistakes, and move confidently toward publication.
Frequently Asked Questions
Which of these mistakes is the most expensive to fix late?
Missing confounders and an unspecified primary outcome are usually the most expensive, since both can require additional data collection — sometimes impossible if the study window has closed.
Is it too late to fix these issues if data collection has already started?
Not necessarily. A data dictionary can still be created retroactively to standardize what's been collected so far, and a primary outcome can often still be specified before analysis begins — the earlier, the better, but partway through is not the same as after.
Do these mistakes apply to retrospective studies too?
Yes, arguably more so. Retrospective designs make it easy to skip a formal protocol since the data already exists, which is exactly when a missing confounder or unclear outcome definition is most likely to go unnoticed until analysis.
How can I catch these mistakes without a formal mentor?
A structured checklist reviewed before data collection begins, plus a second set of eyes — even a colleague outside your specific project — can catch most of what a formal mentor would.