During the acute phase of the COVID pandemic, Nextstrain became known as a key platform for tracking the evolution of SARS-CoV-2. It took viral genome sequences uploaded to global repositories like GISAID and GenBank and turned them into phylogenetic trees that put viral evolution on display.
Now, Nextstrain is moving into what its leaders describe as its next phase, seeking to empower public health professionals to use its tools for their own analyses. Such resources could come in handy as public health professionals across the country battle ongoing measles outbreaks, and as they monitor global outbreaks such as Ebola.
MedPage Today spoke with two leaders of Nextstrain, based at Fred Hutch Cancer Center in Seattle, to better understand the platform and how it can aid public health professionals. J.T. McCrone, PhD, a microbiologist and immunologist specializing in virus evolution, and John Huddleston, PhD, a cellular and molecular biologist, both have substantial biostatistics, bioinformatics, and epidemiology experience.
An edited transcript of the conversation follows.
What is Nextstrain?
McCrone: Nextstrain is a genomic epidemiology platform where phylogenetic analyses of virus evolution and transmission are hosted online. The hole that Nextstrain fills is that a lot of these analyses are done by academic labs around the world. They're done by public health labs. They're done as sort of one-off analyses that are required when there's an outbreak.
Nextstrain has built a number of tools and pipelines that standardize that process, make it reproducible, and then host it on the web so that people can get real-time insights into viral spread and evolution.
While in the past you would wait for an outbreak response report or a publication to update you on what was going on, Nextstrain allows real insights at the speed at which data is being generated.
How does Nextstrain get its data?
McCrone: Viral researchers and genomic epidemiologists generate data around the globe and share that data publicly through a number of platforms. We're talking about genomic sequences of viruses and maybe some metadata on location or where they're found.
Nextstrain uses that publicly available data and analyzes it, and then makes those analyses accessible for others. We're sort of at the mercy of data generators and their goodwill and sharers around the world. Hopefully what people get from Nextstrain is an insight into their data.
All of the analyses that are run by Nextstrain are made publicly available. We have collaborations with public health departments who are running their own Nextstrain builds and hosting them and sharing them externally and internally.
Does Nextstrain work with the CDC?
Huddleston: We collaborate with the CDC, specifically the influenza division. Historically, we used to do analyses where we would construct phylogenetic trees showing relationships of flu viruses, and we would share those analyses with folks there. But since the pandemic, the capacity for them to do their own analyses has really expanded, so they're often at the cutting edge. It's much more of a fruitful mutual experience.
There are many different divisions within the CDC that take advantage of Nextstrain tools. So, in addition to us hosting analyses on Nextstrain, people can create their own versions of that Nextstrain website to use internally. We also provide ways for people to share their data out through a Nextstrain groups interface at nextstrain.org/groups.
How can the public health community use Nextstrain?
McCrone: I think Nextstrain users fit into a couple of categories. One are people who could benefit just from getting a live look at the patterns of viral circulation and evolution. For those individuals, the website is sufficient.
But then individuals at public health departments often will have data they cannot make public immediately, or they'll have data that they want to do an internal analysis on. For those power users who are bioinformatically savvy, all of the pipelines are available and built in such a way that you can augment a publicly available dataset with your own internal data to generate an upstream build to explore and share internally. Those are users that I think we serve fairly well.
There's a middle class of users. Genomic epidemiology is interdisciplinary, and so people come with a lot of different expertise. One of the directions that we're hoping to move in is toward these middle users -- individuals who might not have the bioinformatic resources or expertise to do their own analyses on the command line, but have data that they would like insight into. We're looking at ways to make analyses more accessible for those individuals.
Huddleston: We love hearing from users, and the more people we get to talk to, the better. They can email us at hello@nextstrain.org, and that goes to the whole team. We have a discussion board, discussion.nextstrain.org, where people can post any questions they have. We're hoping that not only is there a community that builds around that public interface, but also we get to hear more of what people want to use these tools for.
There are several viral outbreaks currently: measles and Ebola, and hantavirus ended earlier this summer. How was Nextstrain involved in those?
McCrone: There are builds for all of those viruses on Nextstrain. We've been involved at varying degrees. We collaborate closely with some groups in Europe who were involved with sequencing the hantavirus genomes.
We're in contact with researchers at the [National Institute of Biomedical Research] in the Democratic Republic of Congo for the current Ebola outbreak. There's a lot of support for the researchers there. They obviously have a lot of expertise in tracking Ebola genomic technology. We're in communication with them through our current funding with the Gates Foundation to provide support as needed.
I know collaborators at the Washington Department of Health have used it to monitor the measles situation both in Washington and around the country.
Huddleston: One way that the Nextstrain measles tree is useful for those types of groups is to provide a reference for what is happening now. They might have their own outbreak sequences that haven't been publicly released yet, but they need to know how to contextualize those sequences in the broader genetic relationships of what has circulated in the past. The Nextstrain tree as a reference gives them a place to put those recent data and see what genotypes they belong to, or what genetic groups they belong to.
What were some of Nextstrain's key contributions during the COVID pandemic?
McCrone: I was not part of the Nextstrain team during the COVID pandemic, but as an external individual in the field, it was a fantastic resource to go and look for mutations that might be of interest in their diversity and how rapidly they were spreading. All of the genomes that are hosted on Nextstrain are available to the public in some way or another, and so these are analyses that you could run theoretically, but that takes a lot of effort. To have it as a resource that you could build very quickly to look at a mutation of interest or a variant of interest was a fantastic resource.
Huddleston: I think some of the lessons we learned were more practical computational lessons, like how do you scale up to millions and millions of genomes when SARS-CoV-2 was the most heavily sequenced virus that we'd ever experienced. That really led to whole new approaches, like the Nextclade tool that allows you to drag and drop your sequences onto a browser. It does a reference-based alignment very rapidly and tells you, this is exactly the clade that the sequence belongs to.
We now create Nextstrain reference trees for every pathogen as just a matter of course, based on the work that happened with SARS-CoV-2. That's really helped reach that other audience where maybe they don't have a lot of bioinformatics experience to work on the command line, but they know how to work with these tools in the browser.
The other thing that we've learned is just the power of having modular tools that can easily pull in data from other sources. So we can integrate if someone's doing experimental work in an academic lab that produces a lot of interesting data. How can we easily pull that into our trees to help people visualize it?
Is there anything else you want potential Nextstrain users to know?
McCrone: I think John's comment of hearing from users is really important. One of the main challenges with developing any kind of software, probably academic software in particular, is knowing your user base. We know collaborators who reach out to us, and we work closely with them. But Nextstrain has a very wide audience. So hearing from people about things that they would like to see on the website, about tools they would like to see integrated to use on their own data, is always something we're interested in.
https://www.medpagetoday.com/special-reports/features/122632
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.