Have a request for an upcoming news/science story? Submit a Request

Anvil lowers barrier to entry for big data cancer researchers at BigCARE 2026

  • Science Highlights
  • Anvil

Over the summer, Purdue University’s Anvil supercomputer supported the 2026 BigCARE Workshop, a two-week course that teaches cancer researchers big data skills. The workshop took place at the University of California, Irvine (UCI) Joe C. Wen School of Population & Public Health. Throughout the course, attendees learned to manage, visualize, analyze, and integrate a variety of omics data in cancer studies. Anvil was integral to the workshop, providing attendees with access to a high-performance computing (HPC) resource designed to have a low barrier to entry for newcomers, which is crucial for cancer researchers who may not yet be experts in research computing.

The Big Data Training for Group photo of BigCARE participants and instructorsCancer Research (BigCARE) workshop is a program funded by the National Cancer Institute (NCI). It was founded in 2020 by Min Zhang, MD, PhD, a Professor of Epidemiology and Biostatistics at Wen Public Health and the Biostatistics Shared Resources Director for the UCI Chao Family Comprehensive Cancer Center. Her founding collaborators include: Dr. Sean Davis, MD, PhD, Associate Director of Informatics and Data Science, Professor of Medicine, from the University of Colorado Anschutz School of Medicine, and Dr. Dabao Zhang, PhD, Professor of Epidemiology and Biostatistics at Wen Public Health. Recognizing the need for specialized HPC and Big Data training for cancer researchers, the team designed BigCARE to provide for that need. This year’s workshop focused on teaching 1) the basics of R programming; 2) data analysis workflows for multiple cancer data types, including targeted and untargeted metabolomics data, bulk and single-cell RNA-Seq data, ChIP-seq and ATAC-seq epigenomics data, and microbiome data; and 3) the basic concepts for conducting gene association studies, genotype-instrumented causal inference, and causal gene regulatory networks construction.

“Anvil has been extremely helpful during the previous BigCARE workshops,” says Zhang, “especially for our participants with limited computing skills. Anvil provides the essential infrastructure and computing support needed to navigate between command line and R packages for large-scale data. This year, Anvil made the implementation much smoother when we added some AI and machine learning tools for multi-omics data analysis. The Anvil platform, along with Jupyter Notebook, offered an all-in-one solution that helped participants easily and quickly switch from concept to interactive analysis of big data without obstacles.”

Anvil’s role in the BigCARE workshop was to provide HPC resources through Open OnDemand and Jupyter Notebooks, which limited the need for in-depth knowledge of command-line interfaces or HPC server environments. The course material was developed as Jupyter notebooks, which researchers could access directly via the web thanks to Open OnDemand. All of this equated to a low barrier of entry for the workshop participants.

For BigCARE 2026, staff members from the Rosen Center for Advanced Computing (RCAC) pre-installed all the bioinformatics tools and datasets that participants and instructors used for the workshop, resulting in a smoother workshop experience for all. Eric Adams, the Lead Research Operations Administrator for Education, and Ryan DeRue, a Senior Computational Scientist, also attended the event at UCI to serve as instructors for the HPC/Linux portion of the workshop and provided on-site support.

“This is my third year serving as an instructor and providing on-site support for the BigCARE workshop, and each year I've attended, I am always struck by how impressive the attendees of the workshop are,” says DeRue. “They are comprised of faculty, med students, postdocs, and staff, and each of them is extremely knowledgeable in their focus area within cancer research. So, it's especially fulfilling to me to play a part in connecting their ideas to a system like Anvil where those ideas can culminate. Many of the attendees go on to become computational researchers in their own right, and so I've always viewed the BigCARE workshop as mutually beneficial for the attendees as well as for us. In some sense, we get to cultivate the computational intuition of the researchers that will one day be applying for their own allocation on our systems.”

As in years prior, the BigCARE 2026 workshop was a great success. Dr. Zhang and the attendees were excited by what they accomplished during the two-week intensive, and pleased with how helpful the RCAC support team was. In a post-course survey, 17 participants said they were likely or very likely to apply for their own Anvil allocation in the future, with 27 planning to continue using the system for the remainder of their one-year allocation, which was included as part of the workshop. Dr. Zhang also indicated that she intends to continue using Anvil to support training activities for the foreseeable future.

More information about the BigCARE 2026 Summer Workshop can be found on UCI’s “Big Data Training for Cancer Research” webpage. Information about the Anvil supercomputer can be found on Purdue’s Anvil Website.

For more information regarding HPC and how it can help you, please visit our “Why HPC?” page. Anvil is funded under NSF award No. 2005632. Researchers may request access to Anvil via the ACCESS allocations process or through the NAIRR allocations process.

Image of the Anvil supercomputer inside Purdue's Data Center, including the Anvil AI partition

Written by: Jonathan Poole, poole43@purdue.edu

Originally posted: