NEWS

Does AI Remember Science Fairly? Yixuan Liu Investigates through a Network Science PhD 

        Yixuan Liu came to network science from a diversified background. She studied statistics as an undergraduate and applied math for her master’s, but it was a research stint studying social networks that pulled her toward a new way of thinking about the world: understanding people and systems through their connections, rather than as isolated data points.

        When she began looking at PhD programs, Northeastern’s Network Science Institute stood out immediately, “It’s one of the few places with a whole PhD program built around network science, with a really strong, close-knit community of researchers all approaching complex systems from different angles.”  That same cross-disciplinary spirit now runs through her own research, which sits at the intersection of Artificial Intelligence (AI), science, and fairness. 

Auditing AI’s Memory 

        At the Barabási Lab, Liu studies scientific recognition in AI, meaning how well Large Language Models (LLMs) actually know who scientists are, and how that knowledge is shaped by factors like citation visibility, gender, career stage, and geographic region. LLMs, which are the basis of generative AI tools, such as Anthropic’s Claude and OpenAI’s ChatGPT, are a type of artificial intelligence program that has been trained on vast amounts of text to understand, summarize, translate, and generate human-language responses. Liu’s work spans several proprietary and open-weight models, and the pattern that emerges is consistent: recognition is sparse overall, and heavily skewed toward scientists who are already visible, senior, and Western, even after accounting for actual scientific impact.  

        Using an open-training-data model, Liu and her collaborators show that a model’s “recognition” of a scientist tracks how often their name appears in training text far more than it tracks real scholarly influence. She uses interpretability tools like causal tracing to reveal where a scientist’s recognition physically lives inside a model. Her findings show that recognized scientists get a clean, localized “memory trace,” while unrecognized ones don’t, which also explains why models tend to hallucinate when asked about someone they don’t actually know, making up journal and co-author names and publication dates. 

        The results of her research have implications for real life because these recognition patterns bleed into real scientific practice. Papers written with LLM assistance tend to cite the scientists these models already recognize more than equally accomplished scientists they don’t, “The bias risks reinforcing itself over time, rather than correcting existing inequality in who gets seen and cited in science.” One statistic in particular surprised her, “Even highly cited scientists only get recognized around one in ten times, and a pattern that holds pretty consistently across different AI models.” What is even more surprising is that this ratio doesn’t change as models get newer and more capable, proving the stubborn persistence of this challenge to close the gap between real impact and having AI systems recognize them. 

Why Does This Matter? 

        For Liu, the stakes go well beyond a single research niche, “More and more people are using AI as their first stop for learning about science, such as asking a chatbot who the experts are in a certain field or using AI to help draft literature reviews and research summaries.” If a scientist isn’t well represented in these models, they risk being overlooked regardless of the quality of their work, potentially repeating and automating the same gatekeeping that already exists in science.

        The question extends further still, into how memory and retrieval function in generative AI more broadly, “These models are becoming a new layer of collective memory, and unlike libraries or citation records, that memory isn’t transparent or easy to audit. We can’t just look up why something is remembered or forgotten.”  These gaps are already shaping how the world learns about science; her research brings them into view, so we can guide these tools consciously rather than quietly inherit their blind spots. 

From Individual Scientists to Human-AI Collaboration 

Professor Albert Laszlo Barabasi with Yixuan Liu during a previous annual review

        Liu describes the most exciting part of her PhD as the shift in scale her thinking has undergone, “The most exciting part has been working with people whose papers I used to read as an undergrad, and realizing research is way more fluid and collaborative than it looks from outside.” What began as a focus on individual scientists has expanded into a fascination with human-AI collaboration itself, and how it’s reshaping the practice of science.

        After graduating, Liu wants to keep working at the intersection of AI and the practice of science, studying not just individual tasks like writing or searching, but the deeper structures of how knowledge is produced, credited, and passed on, “As agentic systems and long-context models take on more of the actual work of research – the reading, synthesizing, deciding who to cite or build on – whose contributions these systems recognize becomes inseparable from what gets discovered and pursued next.” Her goal is to help define what a more accountable model of scientific memory could look like, one where visibility and retrieval are traceable and correctable, rather than an opaque byproduct of training data. 

Advice for Future Students 

        Liu’s advice to prospective PhD students is to resist the urge to arrive with a finished plan:  

“A PhD isn’t just about the title at the end, it’s about figuring out what you’re actually curious about, and giving yourself room to explore before narrowing down. Don’t restrict yourself too early. Come in open-minded and okay with not having everything figured out. The best projects here usually don’t start with a clear shape, they start messy and get better the more you dig in and talk to people.” 

        She also credits conversation itself as a research tool: 

“Don’t be afraid to reach out. Some of my best ideas came from a random conversation. Northeastern, especially NetSI [Network Science Institute], is a great place for exactly that kind of exploration. It’s small enough that those conversations actually happen, but full of people working on really different problems.” 

        In reflecting on her time at Northeastern, Liu spoke of a sense of gratitude, feeling “lucky to be here as a NetSIer.” She credits the community she has cultivated at Northeastern – her supervisor Professor Albert-László Barabási, colleagues, collaborators, and friends – as having helped and supported her along her PhD journey. 

 

Photo Credits: Nicole Samay, Dr. Albert Laszlo Barabasi 

 

Yixuan Liu is a PhD candidate in Network Science at Northeastern University, working in the Barabási Lab. Connect with her on LinkedIn or learn more about her work at yxliu.com.