Search and Research
Literature Searching to Support Research Rigor
Lesson 1:
Orienting to a new adventure
Summary:
In this lesson, we consider the foundational purpose of literature search: the acquisition of knowledge. Facing a new subject area, we may feel that we need to learn everything. But recognizing the impossibility of that, we may fail to learn enough, putting our project at risk. Here, we break down the kinds of knowledge we need to safely proceed, set goals to know when we are finished, and chart a course to efficiently find and learn from the resources we need to accomplish them.
Goal:
Set goals and plan a process for developing a knowledge base in a new subject area.
1.1 What are we searching for?
This unit is designed to help make the most of a literature search. But before we get to that, let’s take a moment to recognize: we don’t search the literature for its own sake! Literature search is part of a learning process that ebbs and flows throughout the research cycle. In that broader process, we build our understanding and awareness in ways that are beneficial (always) and critical (often) to the quality of our research and reporting. Literature search is a key element of the research process, but to understand its role we need to work through the process as a whole.
We can’t learn anything from literature we haven’t found, which is why solid search skills are vital. But what else is involved? We can’t benefit from search results unless we understand the importance of what we have found. This demands that we build knowledge before and during a search, set goals and monitor learning, and then refine our search to prioritize resources, making the best use of limited reading time and learning the things we MOST need to know.
So, to tell the story of search, we will start at the beginning: building the knowledge we need to search effectively.
1.2 A first set of goals
Let’s imagine that we are commencing research on a topic that is entirely new to us. For the purposes of this unit, we’ll imagine that we are joining an already active project, or perhaps following someone’s recommendation. It might technically be possible for us to proceed without much exploration.
To illustrate how this might apply, let’s advance our imagined scenario to consider a simplified topic example. The study we are preparing to pursue investigates a relationship between sleep deprivation and memory.
The project
Hypothesis: Sleep deprivation interferes with memory.
Prediction: If we deprive subjects of sleep, their performance on a memory test will decline.
Baseline: Recruit student volunteers; collect demographic data.
Primary outcome measure: Word-list task
Day 1, study a word list
Day 2, measure the number of words recalled.
Treatment group: Total sleep deprivation.
Control group: Standard sleep.
(Note: this scenario doesn’t match most projects in real life! We start here in order to focus on the process of learning and searching. For advice on developing a project idea, be sure to have a look for our unit on research questions!)
We know we’ll need to learn more at some point. But at the outset, haste is not only appealing - it often seems to be demanded at every turn. The pressure to get going can be quite intense. This sets up a hard reality: there is never enough time to learn everything about a project at the outset. This is simply a fact, not a failing. But how we respond makes a difference.
We are about to invest time from our life and career in this project. Precious, hard-won funds are being spent. If we can find a way to efficiently and effectively gain the essential knowledge we need, we can protect that investment and enjoy confidence in our work. If we choose to proceed without first exploring our topic in key ways, that choice exposes us to risk.
What kind of risks are we talking about?
Risks of planning the project wrong:
Say we launch into a project with a method planned, but we haven’t explored the alternatives. Supplies are ordered, arrangements made. Chatting with another scientist, we excitedly relate our plans. They ask: “Wait, why would you use that method?”
Oops: this is not only a moment of embarrassment (a risk in its own right) - we should be wondering if they have a point! Did we even know that there was an alternative? What if it’s a better fit to the question? This could be a very expensive mistake to correct!
Risks of completing the wrong project:
Worse, fast forward to the end of a project. Everything went smoothly (luckily, we had great advice). We take our results to a conference. After we explain our findings, someone asks, “Didn’t so-and-so publish something similar [before the start of the project]?”
This scenario occurs more often than you might think! Studies show that major trials are often missing applicable reviews in their citations (Lund et al. 2022) and that a significant amount of biomedical research wastefully overlaps with prior work (Engelking et al. 2018, Rosengaard et al. 2024). Again, leaving the cringe-inducing nature of this moment aside, the costs of time, money, and career momentum associated with this kind of error are intolerable. And avoidable!
These are not the only risks involved in starting and completing a project without learning key information. However, these two examples are useful in highlighting two types of goals that we can set to minimize our risks without setting ourselves up with the impossible task of learning everything.
1.3 Necessary and sufficient knowledge
We accept that we can’t know everything, and that time is absolutely of the essence in advancing a project. Nonetheless, we have also established that there are certain kinds of information we need to be sure we have learned before investing time and money. This information falls into two categories:
- Knowledge to explain key decisions about our project.
- Confidence that we haven’t missed key publications that could affect our work.
Critically, this helps us begin to set goals for the learning we need to do. The better we can define “enough” in a reasonable and achievable way, the more effectively we can use our learning time to reduce the risk of mistakes. With solid goals and an efficient strategy, a little learning can go a very long way!
Reflection:
- With these broad goals in mind, how would you begin?
- How would you know when you were 'done'?
1.4 Learning is an iterative process
In fact, learning is rarely a linear sequence of steps, but an iterative process. In order to be effective, we want to be intentional about how to seek out information and build our knowledge around a research project.
First, we might identify unknowns in a scientific theory. For a new topic, this could include so many things related to the components of the project or decisions around methodology!
Second, once we have some idea about the unknowns, we want to formulate questions about the unknowns. These questions can include the types of inquiries in the above section, such as "why use this method" or "what is the importance of control X in the project".
Third, we want to take intentional steps to gather information to resolve the gaps in our knowledge. To do this, we'll want to be efficient in our selection and usage of tools and resources, including experts and the published literature.
Lastly, we'll want to integrate the new knowledge back into our existing understanding. This may clear up everything we were curious about previously. It might also reveal new questions about aspects we previously didn't even realize were there!
Let's take some time to discuss where we might obtain information in our learning process. The choices we make can be very personal.
Social and well-connected? You might jump straight to finding a friend with the expertise to get you going in the right direction. A book-worm with confidence in your reading superpowers? You’re probably out searching for review articles and book chapters to get your bearings. Tech savvy or just living in the now? For a lot of people, this journey will begin as a long talk with a favorite LLM.
All of these are extremely useful options! Each of them has pros and cons. We’ll address all of them at different points in this unit. No matter where we look for a first step, at some point early in our process we are likely to reach for a tool: a search engine, a database, or an LLM. In most cases, these are tools we use regularly. But how much do we really know about their strengths and limitations?
Activity #1: Choose the Right Research Tool
First steps often involve a tool. How do you choose? Test your knowledge of the strengths and limitations of databases, search engines, and LLMs.
Post-activity questions:
- Did anything surprise you?
- Are there any other important distinctions you can think of between databases, search engines, and LLMs?
- In what ways might these differences be important in determining which tools to use for certain tasks?
- What mistakes do you think are most common in the choices researchers make?
Increasingly, tools are combined to leverage multiple strengths in response to a single search entry. While this can improve results, it can also make it harder to know exactly what the strengths and limitations of any given search approach are. Still, knowing what questions to ask about a tool is a solid start to get the resources we need.
But what resources are those, exactly? And how will we know when we have what we need? Even when we know where to start, we don’t want to begin without knowing where we should end. How can we expand upon our goals to recognize success in achieving them?
1.5 Setting goals for knowledge
Our first broad goal was “Knowledge to understand and be sure we agree with key decisions about our project.” This is a little less comprehensive than “know everything” but it’s still sufficiently indefinite that it might be hard to know when we’re there. What does this look like?
Let’s revisit the first question we couldn’t answer above: “Why would you use that method?” In order to explain that well, we would need at least two kinds of information.
- We would want to know that the method is a practical and reliable way to produce evidence.
- We would want to know that that particular type of evidence will usefully inform the theory under investigation.
Most new information may be classified as theory, method, or evidence. However, in most cases, that information is meaningless for scientific purposes until it is connected to one or more of the other categories. Right now, we don't have these connections yet, but we will take the time to explore this framework in lesson 2.
Takeaways:
- Literature search is one important part of a broader learning process.
- Two kinds of learning are essential to support a project’s rigor: knowledge to explain key decisions and confidence that we haven’t missed key publications relevant to our project.
We can evaluate the completeness of our understanding of a research concept based on how well we can connect it to theory, method, and evidence. - Applying a goal-directed process for exploratory learning and search helps us avoid overwhelm and accomplish what we need in the time that we have.
Reflection:
- Have you ever put off learning about a topic longer than you intended to? What do you think got in the way?
- If you use structured databases regularly for search, what are the circumstances that lead you to that choice? If you don’t, how do you usually choose a search tool for academic literature?
- Consider an element of the work that you do. Can you draw a complete connection between theory, method, and evidence for that concept? Did you find any gaps?
Lesson 2:
Connecting theory, method, and evidence
Summary:
A key challenge in exploratory learning is to make sense of new knowledge. This lesson introduces the Theory Method Evidence framework for organizing knowledge. The lessons works through an example to identify how seemingly-contradictory studies are actually testing different and more specific claims to refine a shared scientific theory.
Goal:
Connect theory, method, and evidence to deepen understanding of a new research topic; map a broad theory into specific claims that individual studies can support or refute.
2.1 A framework for scientific knowledge
To make sense of a new topic, it is helpful to have a framework for organizing information.
The framework we are using in this unit organizes scientific knowledge into three categories:
Theory
How we think things work, and their alternatives
Method
Available techniques and their challenges.
Evidence
Observations that inform theory or method.
Let's look at how this might apply for our sleep deprivation project.
The primary theory in this study is that sleep deprivation interferes with memory. The methods as developed are to test memory after sleep deprivation; and to compare the performance of human subjects who experienced sleep deprivation and a control group who did not. The evidence that might result from this study is the identification of a possible performance difference between groups.
But theory, method, and evidence are intertwined, so really we need to explore the relationship between these bits of scientific knowledge
1. How Evidence influences Theory
Whether the results do show a difference is information that provides either support or change to our existing theory. We update the theory in light of new evidence.
2. Without theory, methods and evidence are isolated facts.
If we consider just the data collected and the methods used, we know that one group might have different performance on a memory test than the other group. But it is the theory that provides an explanation that we can generalize to other populations or other contexts (depending on the strength of the evidence for the theory.)
3. We need evidence to know that a method is working.
The evidence of a difference between groups is a signal that the sleep deprivation was effective. If we didn't see a difference, it could mean that sleep deprivation had no effect; or it could mean that our method of inducing sleep deprivation wasn't strong enough to influence memory performance. The usage of proper controls comes into play, and more details can be found in our unit on Controls.
4. Proper interpretation of the evidence depends on both the method and the motivating theory.
What the evidence means for other populations or other context depends on the theory and accumulation of evidence from other studies; as well as our confidence that the methods operated as expected (which might also be the result of evidence from prior studies).
When we connect theory, method, and evidence, we can justify the role of project components.
This gives us the knowledge to explain key decisions about our project!
But we still need a place to start…
2.2 Beginning your exploratory learning
Common sources for initial readings are:
- recommendation from an expert/colleague
- results from a casual search
- references of papers you already know
- a recent paper/talk/poster
Review articles are often an effective starting point. They can provide a survey of the topic; provide organization for key ideas and evidence; and have references for key pieces of evidence, enabling us to trace the history of ideas.
However, be aware that reviews aren't neutral and are based on the perspective of the authors. Even systematic reviews and evidence synthesis papers, which are aimed at being comprehensive collections of scientific knowledge in a precisely defined area, are designed by humans who draw the boundaries of what is and is not included. Furthermore, older reviews may be out of date, and will not include newer evidence that changes or refutes extant theories at the time of review.
2.3 Sleep deprivation and memory
Imagine we begin our reading about sleep deprivation and memory…
First, we read a classic of the field, "Effect of REM sleep deprivation on learning and recall by humans" (Chernik, 1972). Let's start by summarizing the knowledge of this paper:
Theory
Sleep deprivation impairs memory.
Method
Human subjects were trained on a verbal recall task. 16 of the subjects were deprived of rapid eye movement (REM) sleep for 2 nights. The other 16 subjects were controls and slept at home.
Evidence:
No significant difference between the two groups.
Next, we read "Effects of Early and Late Nocturnal Sleep on Declarative and Procedural Memory" (Plihal and Born, 1997).
Theory
Sleep deprivation impairs memory.
Method
Human subjects learned a verbal recall task, then were tested after a retention interval. One group slept, another stayed awake. Within each group, the tested intervals included both early sleep (mostly deep, non-REM) and late sleep (mostly REM).
Evidence:
Subjects in the "sleep" condition performed better on the recall test compared to subjects in the "wake" condition.
A contradiction appears
Initially, it seems as though the two papers present conflicting evidence about the effects of sleep deprivation on memory. And there are many more papers in this area that likely also diverge in their methods and interpretation. How can we go about resolving this?
One aspect to keep in mind is that science is a community effort.
Each individual study contributes a tiny bit of evidence to expand on and specify a
scientific theory. A solution that comes to mind then, is to map a broad theory, such as "sleep deprivation impairs memory", into specific claims or predictions that can be confirmed or refuted by individual studies.
For example, we could start by investigating what kind of sleep matters.
Claim 1: The memory benefit of sleep depends on REM sleep.
To clarify how individual studies relate to claims, we can use an "if … then …" construction to produce a prediction. For this claim, the prediction would look something like "if subjects are deprived of REM sleep, they will perform worse on memory tests."
What is the evidence of Chernik (1972) about this claim? Well, in that study, subjects deprived of REM sleep performed just as well as control subjects. So these results are inconsistent with the prediction.
And what about Plihal and Born (1997)? In that study, the benefit of sleep was strongest for early sleep, which has the least REM. So these results are also inconsistent with the prediction.
Claim 2: The memory benefit of sleep depends on deep, early (non-REM) sleep.
For claim 2, the corresponding prediction is "if subjects are deprived of deep, early (non-REM) sleep, they will perform worse on memory tests"
In Chernik (1972), subjects were deprived primarily of REM sleep and performed just as well as control subjects. So these results are consistent with the prediction.
In Plihal and Born (1997), the benefit of sleep was strongest for early sleep, which has the most deep, non-REM sleep. So these results are also consistent with the prediction.
Synthesizing the evidence
By comparing the evidence in these two papers against specific predictions, we found that there was support for claim 2, but not for claim 1. (Though further studies that provide tighter controls on REM vs non-REM sleep could help distinguish this more clearly.)
Activity #2: Evaluate Claims from the Literature
Now it's your turn to practice. We've generated a series of claims for you to evaluate against evidence from some newer studies.
Post-activity questions:
- How did connecting methods to evidence to theory help you identify gaps and outstanding questions?
- Have you ever expressed predictions from scientific theories before? Was it useful?
- For proposed claims, how many studies would it take to support or refute them?
2.4 Rigor matters
Note that our synthesis so far of the findings from multiple studies treated the evidence from those studies as fact. But that depends on how rigorous each study actually was! Remember that the evidence in a study is only as strong as the methods and design. Are the controls appropriate? Is the sample size sufficient? These details can change how confident you are about the evidence and can matter a lot!
Some details of the study that affect rigor:
- What kind of sleep deprivation is happening?
- How well do the methods match the question?
- How is memory being tested?
- What is the sample size?
- What population is being tested?
- How are other factors controlled for, or tested?
- What other factors are there?
- Is the statistical analysis appropriate?
For a more comprehensive look at these details, you may wish to check out these other units:
Controls: Clarity not coincidence!
and coming soon:
Sample Size
Assessing the Rigor of Published Literature
2.5 From a map to its gaps
Reading is necessary but not sufficient for learning. Recall that our goal is to become situated about the topic of sleep deprivation and memory in order to justify components in a new research project. Thoughtful engagement with what we read is both required and helps save us time. In other words, we need to be reading strategically.
As we sort what we read into theory, method, and evidence, we begin to build a knowledge map of what we know; and also of what we do not yet know. These empty spaces or gaps might be real unknowns in the scientific theory that haven't been explored yet. They could also reflect uncertainty or disagreements between researchers who have been investigating the same topic. A gap could also be due to our own ignorance of what has already been done, or underlying assumptions that experts take for granted..
We might be new to our topic, but we still have prior knowledge, and we are growing it as we read. Drawing on this prior knowledge, we can start to be more directed in our exploratory learning of the topic. A few possible questions might arise:
- What do we know, or suspect, about theory, methods, and evidence?
- What different methods have been applied?
- What other theories have been suggested?
- How has prior evidence addressed questions?
- What gaps remain that lack good evidence? Why?
- How much do we know about what is not yet known?
- Do we have a sense of what we might be missing?
In the next lesson, we'll dig into how to begin exploring these gaps.
Takeaways:
- The theory–method–evidence framework organizes new knowledge: theory is how we think things work and their alternatives, method is the available techniques and their challenges, and evidence is the observations that inform theory or method.
- Making sense of theory, method, and evidence requires exploring how they are connected to each other: evidence updates theory, theory keeps methods and evidence from being isolated facts, evidence tells us a method is working, and interpreting evidence depends on both method and theory.
- A broad theory usually can't be confirmed or refuted by a single study. Identifying specific claims and predictions from a theory ("if … then …") helps us make sense of the evidence from individual studies.
- The quality of evidence from a study is only as strong as the rigor of the methods and design behind it.
Reflection:
- Think of a topic you know well. Can you sort a recent finding into theory, method, and evidence?
- Have you encountered two studies that seemed to contradict each other? Looking back, were they testing the same claim, or different ones?
- When you read a paper, do you take its primary result at face value? What would you focus on first to evaluate the strength of the evidence?
