Empowering Research through the Center for High Throughput Computing

Written by Carmella Whittaker, Photography by Steven Nadeau, Graphic Design by Carmella Whittaker

UW-Madison, a top school in research expenditures, produces groundbreaking research in virtually every building and discipline on campus. Just across Union South, research from across the university and country converge in the basement of the Discovery Building – one home of CHTC.  

Established in 2006, the Center for High Throughput Computing (CHTC) provides computing and data power at no cost to any UW-Madison researcher. It combines computers in local data centers and a team that supports development, infrastructure, outreach, and facilitation to make impossible amounts of data manageable. 

CHTC offers two types of computing resources: high throughput computing (HTC) and high-performance computing (HPC). Computing clusters are often associated with HPC; when a job cannot be done on one laptop, an HPC approach divides the problem over multiple processing cores, where each core needs to communicate with the others. For example, intensive chemistry simulations, where the position of each atom is dependent on every other atom present, need HPC.   

In projects that can be divided into millions of smaller jobs, HTC excels. “High energy physics, astronomy, genomics, and many other scientific domains thrive in high throughput computing; you can slice the problem into small parts, [then] bring all your answers together,” system administrator Joe Bartkowiak explains. This division of work means the computers that can host HTC don’t need excessive specifications. “You can do a lot of what you can do through [HTC] on a high-performance cluster. It’s just kind of like going 25 in a Ferrari,” Bartkowiak continues. 

Most research computed on CHTC’s data center involves computer science, engineering, and mathematics, disciplines that “came up with what is computing from its very inception,” as research computing facilitator Danny Morales puts it.  However, disciplines such as biology have become increasingly quantitative, and the variety of jobs that hit CHTC’s cluster reflects that.  

Morales and other biologists have an increasing need for compute where even local data centers aren’t enough. Research like this is made possible through a CHTC collaboration called the Open Science Pool. Bringing national laboratories and universities together, it forms a distributed network of high throughput computing and works to optimize the transfer of these data sets across the country.

A major challenge in biology is that we’re generating massive datasets, but the training to work with them isn’t built into standard biology education. At one point, I wanted to map where sequences align on a viral genome and predict their cleavage sites—similar to CRISPR. That problem scaled to roughly 50 trillion pairwise comparisons… and promptly brought down my local cluster at Florida International University.

– Danny Morales, Research Computing Facilitator

You can imagine [that the OSPool is] a national version of our high-throughput computing [cluster]. Researchers from any U.S. based institution can then submit a job and that then could possibly land on our nodes, … the computers don’t sit idle, it makes for good reporting, people are happy, and people get research done.

– Amber Lim, Research Computing Facilitator

It’s not quid pro quo. If you are a researcher affiliated with a US academic institution… you don’t have to be part of an organization that’s donating compute resources, you don’t have to be part of an organization that’s paying money… You can leverage the OSPool. In my small worldview, this is a great opportunity for researchers at institutions with fewer computing resources. The truly challenging part is getting the word out.

– Joe Bartkowiak, System Administrator

​“If I didn’t bring it down and took over the entire cluster [at FIU], it would take me about 12 years [to complete my research]. … I was able to do [it] in a semester,” says Morales about the Open Science Pool. In this past year, the OSPool reports supporting over 400 million jobs by almost 100 institutions. 

Even though the physical computers in CHTC’s data center have a reach far beyond campus, the server’s computing power is not the center’s main export. HTCondor, an internationally used software for coordinating HTC systems, was created at UW-Madison 40 years ago and is continually developed by CHTC.  

A newer project, Pelican, shortens the physical distance data travels when it is sent across the country. “Pelican sort of operates under the hood of Condor and the OSPool and all of our many initiatives to make sure that data is actually delivered securely, conveniently, and quickly to wherever it’s needed,” Morales explains. Caches in this network​ copy data that passes through them, functioning as a checkpoint for streamlined data transfer. 

This symbiosis of providing computing for research and the research into computing relies on feedback. Facilitators want to hear about hardware limitations, cache errors, and compute requirements. “We have a couple of users… that constantly push our system to the edges. Part of that trust that we built is, when you as a user say [you] have this issue… I can go down the hall and go to our developers and say this is something that we need to push forward,” Morales states. 

CHTC is constantly making these adjustments and recently focused on troubleshooting a new bioinformatics tool: AlphaFold3, an inference pipeline from Google DeepMind used to predict protein folding. AlphaFold3’s CPU requirements are intensive, and the disk space demands are even greater. “It’s essentially making massive multiple sequence alignments,” Morales says. However, CHTC is developing solutions to these challenges. “

​Rather than the user having to, in the job, download a terabyte of data, which limits the amount of machines that they can hit, we can pre-stage it on a couple of machines, and many more jobs can hit that machine.” 

AlphaFold, sequencing analysis, and complex image analysis are necessary for biological research, and CHTC is making strides to connect this field with the computing resources it needs. This year, CHTC launched the Bioinformatics Cafe. “Bioinformatics Cafe is a pilot program [where] we are testing out and seeing how we want to start interacting with this community that is severely underserved when it comes to computational resources but is in huge necessity of it. We want to start establishing that trust with groups [and start] to build back community,” Morales says. 

More information about CHTC, its services, and how to integrate it with your own research at chtc.cs.wisc.edu. Information about the Bioinformatics Cafe can be found at https://chtc.cs.wisc.edu/uw-research-computing/bioinformatics-cafe

“We never want compute to be the bottleneck [or] reason why discovery isn’t made or people’s lives aren’t changed,” concludes Morales. “We’re trying to empower researchers … in order to advance their lab’s work, their research, and hopefully get students on a path to graduating a little sooner.” ​    

See this article and photo tour of the CHTC data center in our Spring 2026 issue in our GALLERY!

Leave a Reply

Your email address will not be published. Required fields are marked *