The Chatbot Will See You Now: Concerns about Artificial Intelligence and Medicine

Editor’s note: We welcome a guest blog post from Rebecca Kyser, Research and Instruction Librarian at Himmelfarb Health Sciences Library, The George Washington University.

Disclaimer: All opinions expressed in this post are the author’s own and do not represent her place of work.

As a medical librarian who specializes in medical misinformation and disinformation, I tend to have a healthy skepticism for cure-alls. So when I first heard about AI’s promises of automating workflows, eliminating repetitive tasks and curing disease, my first thought was to investigate exactly how true they were. It’s been around four years since then but I’ve emerged from this process to find myself firmly an AI skeptic, especially in regards to large language models (LLMs). And to my disappointment, I’ve found LLMs frequently cited as a solution to solve countless problems of medicine. 

With that in mind, I wanted to share some of my concerns about AI being integrated into medical technology. Not all of these issues are exclusive to medicine – many are found in any field that AI touches – but some of the consequences of AI behaving badly get far worse when people’s health is involved. 

AI is Inaccurate

At this point, almost everyone has heard about AI hallucinations: instances where AI makes up false information. Hallucinations are baked in the technology: there is no way of getting rid of them entirely 1. Rates of false answers vary from model to model 2, and the rate is often published by independent assessors, rather than the companies themselves. The best response to this problem is to lower the rate as much as possible.

However, when we’re talking about wrong answers regarding medical information, what threshold of inaccuracy do we find acceptable? Hallucination rates of 39.6% 2 may be acceptable when asking a bot about the weather, but would we consider that rate acceptable, or even a 10% hallucination rate, for a doctor prescribing you medicine? This isn’t to say humans don’t make medical errors – there’s plenty of scholarship on the topic and how to prevent it – but are their error rates comparable with a bot?

The error rate isn’t great at the moment. Multiple studies have shown chatbots make mistakes regarding medical queries: surgical errors leading to patient injury 3, undertriaging cases 4 and providing faulty medical advice to the public 5. Those mistakes aren’t small outliers either: ChatGPT undertriaged cases 52% of the time 4 and provided correct relevant conditions in less than 34.5% of cases5

There is the argument that as AI gets better (which is an assumption worth investigating), this will cease to be an issue. However, even if AI had an accuracy rate equal or less than humans, another problem arises: who can be held liable for medical error when the error comes from a machine? Is it the person operating the machine or the person who built it? Without these kinds of clearly defined policies, addressing mistakes with AI in medicine becomes a thorny business. 

Garbage in, Garbage Out

Most people know that LLM are trained off broad swaths of information online, but I don’t think we always consider exactly what that entails. Without having access to these models’ training data directly, we can only make educational guesses on what the training data entails, usually by looking at outputs and trying to reverse engineer their source. This is something I’ve done myself: if you have access to an AI that provides sources, ask it about a headline from Alex Jones’ infamous platform Infowars and watch it regurgitate it back to you.

A picture of a headline from Alex Jones Infowars website, showing Terminator-style robots with the headline Humanoid Soldiers Tested in Ukraine; Founder Eyes Contract To Patrol US Border
A picture of a headline from Alex Jones Infowars website, showing Terminator-style robots with the headline “Humanoid Soldiers Tested in Ukraine; Founder Eyes Contract To Patrol US Border”

There is an argument this is remedied by ensuring the model only works off high quality training data. The problem is, we often don’t know exactly what was in the training data to begin with. Even if a model claims to be based on entirely scholarly sources, we should understand all the nuances of what that means: are all the sources peer-reviewed? Does the training set have safeguards against poor-quality sources from predatory publishers? Are retracted publications excluded from the training data? Is the data set monitored for works that are retracted after they enter the training data?

I think the best way to demonstrate this problem is to examine one of the few cases we’ve seen under the hood for a LLM: Claude. Due to a lawsuit, a database was published containing books Claude was trained off of, which the public can access. 

One of the names that comes up in the training set is Graham Hancock. For those who don’t recognize the name, Graham Hancock is a pseudohistorian, who often argues that the work of ancient civilizations was actually built by aliens. His inclusion isn’t surprising – his work is incredibly popular and Claude debunks his theories if asked – but it does raise concerns about the training data for these models. Especially when there are more concerning inclusions, such as the work of Richard Lynn, who can also be found in the data set. 

Richard Lynn was a well known scientific racist, known for promoting pro-eugenic ideas up until his death in 2023. Why is a scientific racist’s work fed into this model?  If this model is being trained to answer questions, why does it have disproven scientific racism in it? You can’t argue its inclusion is due to popularity; unlike Hancock, only one of these publications is well known enough to even have its own Wikipedia page, and that page is mostly about its racist background 6.

I use Lynn as an example because of the history of scientific racism in medicine. Many medical AI attest to training only off medical publications or scholarly journals, but that does not mean those are free of false or harmful ideas. Lynn is an example of this: he published multiple works promoting scientific racism 7. How do we ensure this kind of work doesn’t end up in these models? And is anyone actually looking out for them?

Bias

Many people like to think machines aren’t biased, but that simply isn’t true. Machines are built by humans who contain a wide variety of biases, and those biases get filtered down into both our creations and the data we feed them. A really good example of this is from ProPublica’s investigation of recidivism calculators, which declared Black people at higher risk of offending than white people, even when the white person had a lengthier criminal record. 8 

This is also an issue in medicine. One of the most common mistakes I see in these machines when it comes to medicine is anchoring bias: the bias to assume your first assumption is right. For example, when I asked an AI model for medicine for a differential diagnosis for a patient experiencing chest pain, it suggested cardiac causes. However, when I asked the same question but stated the patient had a history of anxiety, the differential shifted entirely to psychological causes. 

What’s likely happening here has to do with the training data: these models are being fed on content like textbooks, case reports and other educational content. A classic case presentation often gives the reader the information they need to make a diagnosis but might leave out superfluous details that a patient might give in real life. Since the bot is unused to seeing those details outside of cases where they are important, it sometimes anchors on the wrong idea. This may explain why the accuracy rate fell so drastically when it came to doctors inputting information over patients: the doctors knew what information was clinically relevant, the patients did not. 5

No Free Lunch

While many LLMs are free to access at the moment, it seems unlikely that they will remain so. The power and processing power required to run these LLMs is significant and these companies are not non-profits. Which raises a question: how will these companies make money anyway? As the old economic saying goes, there is no free lunch. 

Currently, that’s a question without an answer: OpenAI, the owner of ChatGPT, is still running at a loss and isn’t expected to turn a profit until 2029. 9 They have announced planned revenue streams involving integrating their product into their software (think of an LLM embedded into Word or another consumer product)10, but I want to focus on three potential answers to this revenue problem in particular.

The first revenue option is subscriptions: most AI companies already offer fee-based options such as early access to newer versions of the model or custom-built models with training data specified by the user. Which raises an issue common in subscription models: the cost of access. The inflationary costs of scholarly journals and publications have long posed a problem for library budgets; what happens when we add AI subscriptions to the mix? Will access to the latest AI models widen inequities in education?

Another likely answer to the revenue problem is advertisements. There is a long history of web platforms turning to advertisers to generate revenue when users are reluctant to pay for a service. This might seem more obnoxious than harmful, but the stakes when advertising medical interventions are much higher than those in marketing new shoes. Marketing to doctors fertilized the seeds of the Opioid Crisis, and we should consider the ramifications of Open Evidence- an AI platform that advertises itself to medical professionals- potentially advertising some drugs over others (this idea is explored fully in an excellent post from the Krafty Librarian)

We also have to consider a revenue stream most commonly used by social media sites: the sale of user data to advertisers. Of course, there is a hiccup here: individually identifiable health information is protected under HIPAA and cannot be disclosed. But that doesn’t mean data can’t be used if it is anonymized.

Let’s try an example. In this hypothetical, let’s say there’s a small town in Southern Colorado that suffers from a pollution issue that leads to interstitial respiratory disease. Medical professionals in the area are shown to ask one AI company queries regarding this condition and how to treat it at a higher rate than the rest of the country. The AI Company then sells this data, which while anonymized, shows this community suffers from this issue.

Here are some potential buyers:

  • A supplement company coming out with a tonic to soothe chronic cough. They place advertisements for this area around the symptoms of interstitial respiratory disease and claim it is a “natural option” to treat the condition. Their sales boom.
  • An insurance provider for the area buys the data. Realizing they will be seeing a spike in claims for some treatments before those claims even enter their system, they amend clinics patients can go to in order to cut costs. Less people are able to access health care.
  • A pharmaceutical company takes out advertisements for their latest drug to treat chronic cough. Patients start requesting this drug at local medical providers after being told it’s “the best” despite there being a cheaper generic. 
  • A company that wants to cut corners with their air filters buys land in this area, knowing their additional pollution may not be noticed as quickly.

And so on, and so forth. While companies can already do some of these things by tracking data of searches via Google, the consolidation of demographic information under medical providers would make the task much easier. Are such privacy violations really worth saving a few minutes? 

I don’t wish to cast an entirely bleak view of artificial intelligence in medicine, despite my list above. There are potential use cases with narrow AI; AI trained to do one specific task. However, I recommended a healthy level of skepticism whenever interfacing with an AI product. “Move Fast and Break Things” might be the motto of Silicon Valley, but we cannot let breaking people become a standard practice in medicine.

Acknowledgements: Thank you to Brie McDonald for help editing this piece.

References

1. Nicola Jones. AI hallucinations can’t be stopped — but these techniques can limit their damage. Nature Web site. https://www.nature.com/articles/d41586-025-00068-5. Accessed 3/17, 2026.

2. Chelli M, Descamps J, Lavoué V, et al. Hallucination rates and reference accuracy of ChatGPT and bard for systematic reviews: Comparative analysis. J Med Internet Res. 2024;26:e53164. https://www.jmir.org/2024/1/e53164https://doi.org/10.2196/53164http://www.ncbi.nlm.nih.gov/pubmed/38776130. doi: 10.2196/53164.

3. Jaimi Dowdell, , Steve Stecklow, Chad Terhune, and, Rachael Levy. As AI enters the operating room, reports arise of botched surgeries and misidentified body parts. https://www.reuters.com/investigations/ai-enters-operating-room-reports-arise-botched-surgeries-misidentified-body-2026-02-09/. Updated 2026. Accessed 3/17, 2026.

4. Ramaswamy A, Tyagi A, Hugo H, et al. ChatGPT health performance in a structured test of triage recommendations. Nat Med. 2026. https://doi.org/10.1038/s41591-026-04297-7. doi: 10.1038/s41591-026-04297-7.

5. Bean AM, Payne RE, Parsons G, et al. Reliability of LLMs as medical assistants for the general public: A randomized preregistered study. Nat Med. 2026;32(2):609–615. https://doi.org/10.1038/s41591-025-04074-y. doi: 10.1038/s41591-025-04074-y.

6. IQ and the wealth of nations, Wikipedia Web site. https://en.wikipedia.org/wiki/IQ_and_the_Wealth_of_Nations. Updated 2026. Accessed 3/17, 2026.

7. Dan Samorodnitsky, Kevin Bird, et al. Journals that published richard lynn’s racist ‘research’ articles should retract them. https://www.statnews.com/2024/06/20/richard-lynn-racist-research-articles-journals-retractions/. Updated 2024. Accessed 3/17, 2026.

8. Julia Angwin, Jeff Larson, Surya Mattu and Lauren Kirchner. Machine bias. ProPublica. 2016. https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing.

9. Craig S. Smith. What large models cost you – there is no free AI lunch. Forbes Web site. https://www.forbes.com/sites/craigsmith/2023/09/08/what-large-models-cost-you–there-is-no-free-ai-lunch/. Updated 2023. Accessed 3/17, 2026.

10. Dave Smith. OpenAI says it plans to report stunning annual losses through 2028—and then turn wildly profitable just two years later. Fortune Web site. https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/. Updated 2025.

Amping up Diversity & Inclusivity in Medical Librarianship

This past week, I attended the 118th Medical Library Association (MLA) Annual Meeting in Atlanta, GA. While it was a standard conference in many respects, it was also a historic one. Beverly Murphy was named the first African American president of MLA since it was incepted in 1898.

When I first considered becoming a librarian, I quickly learned about #critlib, which centers the impact of oppression and marginalization of the many –isms in librarianship. I wanted to be in a profession where I could provide information in a critical way, dismantling library neutrality. I found this through a hashtag which allowed me to meet diverse, inspiring, kind, and intelligent librarians. However, I find it slightly more difficult to apply a social justice framework as an academic medical librarian focusing upon the School of Medicine. I have tried my best through critical search strategies and educating others about bias within publishing. And of course, subject areas specific to public and/or global health easily lend themselves to health disparities. Overall though, I have noticed that medical librarianship has been slower to the game, especially in terms of coming together as a community. During this meeting, however, it felt different.

The annual Janet Doe lecture was given by Elaine Martin, focused upon social justice. I have listened to some talks concerning social justice that just scratch the surface. They seem to give a nod to diversity as more of a check box rather than a critical interpretation and call for action. However, Elaine stressed mass incarceration as a public health issue; she emphasized dismantling library neutrality; she quoted Paulo Freire, the author of the seminal Pedagogy of the Oppressed. She received a standing ovation. It was inspiring, and while it may have just been pure emotion, it gave me hope.

I also attended a Diversity & Inclusivity Fishbowl session by MLA’s Diversity and Inclusivity Task Force. During a fishbowl, a moderator poses a question to a group of individuals seated in a few concentric circles. In our case, there were around 30 of us. There were four seats in the innermost circle, and the individuals in that circle answered the question and can be “tapped out” by others in the outside circles who wish to speak. Unless we were in the inner circle, we were solely active listeners. I’m not going to lie, when I saw the format of this meeting, which was three days into the conference and from 5:00 p.m.-6:30 p.m., I dreaded it. But I also knew this was an important issue. Not only did I feel welcome, but I enjoyed the structured yet conversational format. It can be difficult to talk about diversity and inclusion because everyone’s positions are well-intentioned, however, because this is an issue that historically induces trauma upon the marginalized, it can become very passionate. This passion is essential for affecting change, and this format provided a way to combine this passion with respect and compassion. While this is just the beginning of these discussions, it is important to understand perspectives, especially for those greatly affected by oppressions. It was assuring to see so many people coming together while sharing their individual experiences and beliefs for a topic I thought was somewhat dormant within medical librarianship. And, because of the incoming presidency of Beverly Murphy, I am full of hope and faith that events like these will result in an action plan.

I can’t say that I remember everything that Beverly said during the talk she gave after being named the new MLA president. But I can tell you how I felt in response. First, Beverly did not stand at the podium when delivering her words. She sat at a table on the stage to be in conversation with the MLA members. She included song, humor, and love in her words. It was warm. It was inviting. And given the previous events I witnessed, it felt promising. She incorporated the importance of diversity and inclusivity, so it wasn’t a mere check box. Rather, it was always part of the conversation. Just two days before, I met Beverly at the New Members Breakfast. As a co-convener of the MLA New Members Special Interest Group (SIG), I was interested in how we can further engage new members. Shannon Jones, the founder of the New Members SIG, was eager to share ideas with me and introduced me to Beverly, who immediately stated her commitment to advocating for new members. She also told me that she was asking first-time attendees she met to share their experiences, positive and negative, and to contact her directly. Real change comes from strong leaders and action. And diversity is more than an initiative – it is a way of being. Regardless of topic, subject area, or library role, it needs to be part of all we do. Beverly is firm in this commitment:

“No matter what race we are, what color we are, what ethnicity we are, what gender we have, or whether we have physical issues – we are all information professionals, with a common goal, and that is ‘to be an association of the most visible, valued, and trusted health information experts.’ Diversity drives excellence and makes us smarter, especially when we welcome it into our lives, our libraries, and our profession.” – Beverly Murphy

The solidarity and volume is increasing for diverse voices in medical librarianship, becoming a stronger driver for diverse and inclusive representation, pedagogy, scholarship, community, and more and vice versa. I know that equity of race, sexual orientation, gender, and ability is a long road. And I am appreciative we are on it.

 

Narrative as Evidence

This past week I attended the MLGSCA & NCNMLG Joint Meeting in Scottsdale, AZ. What do all these letters mean, you ask? They stand for the Medical Library Group of Southern California and Arizona and Northern California and Nevada Medical Library Group. So basically it was a western regional meeting of medical librarians. I attended sessions covering topics including survey design, information literacy assessment, National Library of Medicine updates, using Python to navigate e-mail reference, systematic reviews, and so many engaging posters! Of course, it was also an excellent opportunity to network with others and learn what different institutions are doing.

The survey design course was especially informative. As we know, surveys are a critical tool used by librarians. I learned how certain question types (ranking, for example) can be misleading, how to avoid asking double-barreled questions, and how to not ask a leading question (i.e. Do you really really love the library?!?) Of course, these survey design practices reduce bias and attempt to represent the most accurate results. The instructor, Deborah Charbonneau, reiterated that you can only do the best you can with surveys. And while this seems obvious, I feel that librarians can be a little perfectionistic. But let’s be real. It’s hard to know exactly what everyone thinks and wants through a survey. So yes, you can only do the best you can.

The posters and presentations about systematic reviews covered evidence-based medicine. As I discussed in my previous post, the evidence-based pyramid prioritizes research that reduces bias. Sackett, Rosenberg, Gray, Haynes, and Richardson (1996) helped to conceptualize the three-legged stool of evidence based practice. Essentially, evidence-based clinical decisions should consider the best of (1) the best research evidence, (2) clinical expertise, and (3) patient values and preferences. As medical librarians we generally focus on delivering strategies for the best research evidence. Simple enough, right? Overall, the conference was informative, social, and not overwhelming – three things I enjoy.

On my flight home, my center shifted from medical librarianship to Joan Didion’s Slouching Towards Bethlehem. The only essay I had previously read in this collection of essays was “On Keeping a Notebook”. I had been assigned this essay for a memoir writing class I took a few years ago. (I promise this is going somewhere.)  In this essay, Didion discusses how she has kept a form of a notebook, not a diary, since she was a child. Within these notebooks were random notes about people or things she saw, heard, and perhaps they included a time/location. These tidbits couldn’t possibly mean anything to anyone else except her. And that was the point. The pieces of information she jotted down over the years gave her reminders of who she was at that time. How she felt.

I took this memoir class in 2015 at Story Studio Chicago, a lofty spot in the Ravenswood neighborhood of Chicago. It was trendy and up and coming. At the time, I had just gotten divorced, my dad had died two years prior, and I discovered my passion for writing at the age of 33. So, I was certainly feeling quite up and coming (and hopefully I was also trendy). Her essay was powerful and resonated with me (as it has for so many others). After I started library school, I slowed down with my personal writing and focused on working and getting my degree, allowing me to land a fantastic job at UCLA! Now that I’m mostly settled in to all the newness, I have renewed my commitment to writing and reading memoir/creative non-fiction. I feel up and coming once again after all these new changes in my life.

As my plane ascended, I opened the book and saw that I had left off right at this essay. I found myself quietly verbalizing “Wow” and “Yeah” multiples times during my flight. I was grateful that the hum of the plane drowned out my voice, but I also didn’t care if anyone heard me. Because if they did, I would tell them why. I would say that the memories we have are really defined by who we were at that time. I would add that memory recall is actually not that reliable. Ultimately, our personal narrative is based upon the scatterplot of our lives: our actual past, present, future; our imagined past, present, future; our fantasized past, present, and future. As Didion (2000) states:

I think we are well advised to keep on nodding terms with the people we used to be, whether we find them attractive company or not. Otherwise they turn up unannounced and surprise us, come hammering on the mind’s door at 4 a.m. of a bad night and demand to know who deserted them, who betrayed them, who is going to make amends. We forget all too soon the things we thought we could never forget. We forget the loves and the betrayals alike, forget what we whispered and what we screamed, forget who we were. (p. 124)

What does this have to do with evidence-based medicine? Well, leaving a medical library conference and floating into this essay felt like polar opposites. But were they? While re-reading this essay, I found myself considering how reducing bias (or increasing perspectives) in research evidence and personal narrative can be connected. They may not seem so, but they are really part of a larger scholarly conversation. While medical librarians focus upon the research aspect of this three-legged stool, we cannot forget that clinical expertise (based upon personal experience) and patient perspective (also based upon personal experience) provide the remaining foundation for this stool.

I also wonder about how our experiences are reflected. Are we remembering who we were when we decided to become librarians? What were our goals? Hopes? Dreams? Look back at that essay you wrote when you applied to school. Look back at a picture of yourself from that time. Who were you? What did you want? Who was annoying you? What were you really yearning to purchase at the time? Did Netflix or Amazon Prime even exist?? Keeping on “nodding terms” with these people allows us to not let these former selves “turn up unannounced”. It allows us to ground ourselves and remember where we came from and how we came to be. And it is a good reminder that our narratives are our personal evidence, and they affect how we perceive and deliver “unbiased” information. I believe that the library is never neutral. So I am always wary to claim a lack of bias with research, no matter what. I prefer to be transparent about the strengths of evidence-based research and its pitfalls.

A couple creative ways I have seen this reflected in medicine is through narrative medicine, JAMA Poetry and Medicine, and Expert Opinions, the bottom of the evidence-based pyramid, in journals. Yes, these are biased. But I think it’s critical that we not forget that medicine ultimately heals the human body which is comprised of the human experience. Greenhalgh and Hurwitz (1999) propose:

At its most arid, modern medicine lacks a metric for existential qualities such as the inner hurt, despair, hope, grief, and moral pain that frequently accompany, and often indeed constitute, the illnesses from which people suffer. The relentless substitution during the course of medical training of skills deemed “scientific”—those that are eminently measurable but unavoidably reductionist—for those that are fundamentally linguistic, empathic, and interpretive should be seen as anything but a successful feature of the modern curriculum. (p. 50)

Medical librarians are not doctors. But librarians are purveyors of stories, so I do think we reside in more legs of this evidence-based stool. I would encourage all types of librarians to seek these outside perspectives to ground themselves in the everyday stories of healthcare professionals, patients, and of ourselves.

 

References

  1. Didion, J. (2000). Slouching towards Bethlehem. New York: Modern Library.
  2. Greenhalgh, T., & Hurwitz, B. (1999). Why study narrative? BMJ: British Medical Journal, 318(7175), 48–50.
  3. Sackett D.L., Rosenberg W.M., Gray J.A., Haynes R.B., & Richardson W.S. (1996). Evidence based medicine: What it is and what it isn’t. BMJ: British Medical Journal, 312(7023), 71–2. doi: 10.1136/bmj.312.7023.71.

 

Questioning the Evidence-Based Pyramid

As a first year health sciences librarian, I have not yet conducted a systematic review. However, as a speech-language pathologist, I learned about evidence-based medicine and the importance of clinical expertise combined with clinical evidence and patient values. As a librarian, I’m now able to combine these experiences, allowing me to view see evidence-based medicine more holistically.

In the past month, I attended two professional development courses. The first was a Systematic Review Workshop held by the University of Pittsburgh. The second was an Edward Tufte course titled “Presenting Data and Information”. While these are two seemingly unrelated subjects, I left both reconsidering how we literally and figuratively view evidence-based medicine.

One of my biggest takeaways from the Systematic Review workshop was that a purpose of  systematic reviews is to search for evidence on a specific topic in order limit bias. This is done by searching multiple databases, reviewing grey literature, and having multiple team members  to screen papers and resolve disputes. One of my biggest takeaways from the Tufte course was that space should be used well to effectively arrange information and that displayed content should have integrity. In his book Visual Explanations, Tufte poses the following questions to test the integrity of information design (p. 70):

  • Is the display revealing the truth?
  • Is the representation accurate?
  • Are the data carefully documented?
  • Do the methods of display avoid spurious readings of the data?
  • Are appropriate comparisons and contexts shown?

When I think about visualization of evidence-based medicine, the evidence-based pyramid immediately comes to mind. It is an image used in many presentations related to evidence-based medicine:

EBM Pyramid and EBM Page Generator, copyright 2006 Trustees of Dartmouth College and Yale University. All Rights Reserved. Produced by Jan Glover, David Izzo, Karen Odato and Lei Wang.

While there is a lot of information in this image, I don’t think it is very clear. I have spoken to librarians (in the health sciences and not in the health sciences) that agree. I think this is a problem. I don’t think all librarians need to immediately know what cohort studies are, but I do think they should understand its context within the visual.

From what I have gathered and discussed with other professionals, quality of evidence/limited bias increases as you go up the pyramid. The pyramid is often explained in a hierarchical way; systematic reviews are considered highest standard of evidence, which is why it is at the top. There are usually fewer systematic reviews (since they take a long time and gather all the available literature about one topic), so the apex also indicates the least quantity. So let’s take a look each of the integrity questions about information design and investigate this further:

Is the display revealing the truth?

Is it? How do we know if this truthfully represent the quantity of each type of study/information? I believe that systematic reviews are probably the least in quantity and expert opinion are the most in quantity. That makes logical sense given the level of difficulty to produce and disperse this type of information. However, what about the types of research in between? Also, is one type of evidence inherently less biased than the ones below? Several studies suggest that systematic reviews may be systematic, but are not always transparent or completely reported and are outdated. This includes systematic reviews published in Cochrane, the highest standard of systematic reviews. While there are standards, they are very frequently not followed. However, following these standards can be very challenging and paradoxical. It’s very possible that a cohort study can be designed in a way that is much more systematic and informed than even a systematic review.

Is the representation accurate?

When I see the word “representation”, I am thinking about visual representation – the pyramid shape itself. There is an assumed hierarchy not just in terms of evidence, but also superiority here. This is a simplistic and elitist way of thinking about this information rather than being informative and useful. If you think about it, a systematic review cannot be conducted without having supporting RCT’s or case reports, etc. Research had to start somewhere. It this was seen as more of a scholarly conversation, I wonder if there would be a place for hierarchy.

I have learned that the slices of the pyramid represent the quantity of publications of each level of evidence. However, this is not something that can be easily understood by looking at this visual alone. Also, if the sizes of the slices represent quantity, why so? Quality is indicated in this version with the arrow going up the pyramid. This helps to represent idea of quality and quantity. However, if evidence-based medicine wants to prioritize quality, maybe the sizes of the slices should represent the quality, not quantity, of evidence. If it is viewed from that perspective, the systematic review slice should be the biggest because it is ideally the highest quality. Or, should the slices represent the amount of bias? This is all quite unclear.

Are the data carefully documented? Do the methods of display avoid spurious readings of the data?

I don’t believe that any data is actually represented here. Moreso, it feels like it’s being told to us so we believe it. I understand this is a visual model, but this image has been floating around so much that it is taken as the truth. I don’t think one can avoid spurious readings of the data because data aren’t represented here.

Are appropriate comparisons and contexts shown?

I do think that this pyramid provides visual way to compare information, however, I don’t think contexts are shown. Again, should the amount of each level of evidence referring quantity or quality? Is the context meant to indicate research superiority? If not, perhaps a pyramid isn’t the best shape. By virtue of its definition, a pyramid has an apex at the top, indicating superiority. Maybe a different shape or representation can provide alternate contexts.

So, how should evidence-based medicine be represented?

I have presented my own perceptions sprinkled with perceptions from others. I’m a new librarian, and my opinion has value. However, I also think this concept needs to be re-envisioned collectively with healthcare practitioners, researchers, librarians, and patients.

Another visualization that has been proposed is the Health Care Literature Wedge. It would look like  a triangle with the apex facing right indicating progressive research stages. I do think there are other shapes or concepts to consider. Perhaps concentric circles? Perhaps this can be a sort of spectrum? 3D maybe? I really don’t know. Another concept to consider is that systematic reviews are intended to reduce bias pertaining to a research question. Instead of reducing bias, maybe we can look at systematic reviews as having increased perspectives? How could this change the way evidence-based medicine is visualized?

I think the questions posed by Tufte can help to guide this. And I’m sure there are other questions and models than can also help. I would love to hear other epistemologies and/or models, so please share!

References

  1. Chang, S. M., Bass, E. B., Berkman, N., Carey, T. S., Kane, R. L., Lau, J., & Ratichek, S. (2013). Challenges in implementing The Institute of Medicine systematic review standards. Systematic Reviews, 2, 69. http://doi.org/10.1186/2046-4053-2-69
  2. Garritty, C., Tsertsvadze, A., Tricco, A. C., Sampson, M., & Moher, D. (2010). Updating Systematic Reviews: An International Survey. PLoS ONE, 5(4), e9914. http://doi.org/10.1371/journal.pone.0009914
  3. IOM (Institute of Medicine). (2011). Finding What Works in Health Care: Standards for Systematic Reviews. Washington, DC: The National Academies Press.) Retrieved from http://www.nationalacademies.org/hmd/Reports/2011/Finding-What-Works-in-Health-Care-Standards-for-Systematic-Reviews.aspx
  4. McKibbon, K. A. (1998). Evidence-based practice. Bulletin of the Medical Library Association, 86(3), 396–401.
  5. The PLoS Medicine Editors. (2007). Many Reviews Are Systematic but Some Are More Transparent and Completely Reported than Others. PLoS Medicine, 4(3), e147. http://doi.org/10.1371/journal.pmed.0040147
  6. Tufte, E. R. (1997). Visual Explanations: Images and Quantities, Evidence and Narrative. Cheshire, CT: Graphics Press.