Research Data Spotlight β Edition 1¶
Spotlight on data by
Niek Barmentlo Β· Ecology & Evolution
Interviewed on 2024-10-28
"I didn't expect to care about FAIR, but it is surprisingly fun and helps me structure my thoughts and project."
-
Keywords
EcologyGenomicsDNA sequencingAfrican Swine FeverWild BoarsYoda -
Information of Niek
The interview¶
Can you introduce yourself and tell us which research group you are part of?¶
I am part of the Ecology & Evolution section, in particular the Evolution group. It is a fairly broad group, but we all tend to use evolutionary theory to explain ecological observations in either controlled or wild settings.
Can you briefly explain your research and what inspired you to work in this area?¶
My project investigates the paradoxical resistance of certain free-living wild boar to the disease caused by African Swine Fever Virus. This pathogen was thought, and observed in controlled situations, to cause 100% lethality in wild boars. However, in wild situations, researchers have observed surviving boar. The question is now what is different between these groups?
I use a combination of genomics, transcriptomics and microbiomics to tackle this question. For me, this research is incredibly exciting as it combines conservation and evolutionary research, implying that I have the "huh, how does this work, evolutionary-speaking?" moments while helping guide conservation efforts!
Have you made your data available yet? If so, where can we find it?¶
Most of my data is currently not downloadable, but all metadata has been written down. For storage purposes, I tend to use Yoda, it is easy enough to use, and allows for both active storage of extra sensitive data as well as 'cheap' archiving of data. Specifically, the archiving is great! When dealing with genomics, one often times only needs the "raw data" for the first step of the bioinformatics pipeline. For me, this means that I require a method to store raw data in an environment that is not my expensive supercomputer storage, and this is where Yoda shines! Additionally, the metadata files needed to be added in Yoda make sure that you can easily match your data to a paper and your code in e.g. a GitHub repository. An example of a small public dataset of mine can be found in Yoda.
What has been the most beneficial aspect of applying data management best practices? Has applying these practices changed the way you approach your research?¶
Genomics datasets are huge! One of my current experiments deals with about 2 terabytes of data, which is sometimes a hassle to keep track of. Especially when you are sequencing samples with companies, you really need to know where your samples came from and what kind of quality and dataset size you can expect from sequencing. Recently, I did a DNA sequencing job with a company in two successions; the first round was to understand how much of the isolated DNA was non-host and the second run was to sequence eligible samples with less contamination in higher quality. The sequencing company messed up the amount of gigabases I wanted per sample, which I only figured out because I knew for each sample what the likely output was supposed to be. All that data management information is stored in large Excels with sufficient README files to keep track of them.
If your data could speak, what advice do you think it would give to other researchers trying to manage their own datasets?¶
In genomics, the data produced is maybe as important as the results produced. If my data could speak, it would say that I should make sure that it is clearly described so it can be used in future studies. As sequencing is expensive, the scientific community should deem it her duty to make sure that labs with less funds can also use genomics data and make sure that scientist don't unnecessarily produce data when it is already available online.
Thank you, Niek!
Thank you, Niek, for sharing your fascinating research and also for sharing thoughtful tips on data management β good luck with your project!