{"id":7777,"date":"2015-10-13T12:00:28","date_gmt":"2015-10-13T11:00:28","guid":{"rendered":"https:\/\/blogs.nature.com\/naturejobs\/?p=7777"},"modified":"2015-10-06T10:24:43","modified_gmt":"2015-10-06T09:24:43","slug":"big-data-collaborative-science","status":"publish","type":"post","link":"https:\/\/blogs.nature.com\/naturejobs\/2015\/10\/13\/big-data-collaborative-science\/","title":{"rendered":"Big data: Collaborative science"},"content":{"rendered":"<h2>The rise of data-intensive research is increasing the need for collaborative science.<\/h2>\n<p>Guest<i> contributor Lakshini Mendis<\/i><\/p>\n<div id=\"attachment_1594\" style=\"width: 310px\" class=\"wp-caption alignright\"><a class=\"wpn-image-link\" href=\"https:\/\/blogs.nature.com\/naturejobs\/files\/2013\/07\/health-data.jpg\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-1594\" class=\"size-medium wp-image-1594 wpn-image\" title=\"data-naturejobs-blog\" alt=\"data-naturejobs-blog\" src=\"https:\/\/blogs.nature.com\/naturejobs\/files\/2013\/07\/health-data-300x233.jpg\" width=\"300\" height=\"233\" srcset=\"https:\/\/blogs.nature.com\/naturejobs\/files\/2013\/07\/health-data-300x233.jpg 300w, https:\/\/blogs.nature.com\/naturejobs\/files\/2013\/07\/health-data.jpg 771w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/a><p id=\"caption-attachment-1594\" class=\"wp-caption-text\">{credit}Image credit: iStock\/Thinkstock{\/credit}<\/p><\/div>\n<p>Big data, a term thought to have <a href=\"https:\/\/www.ssc.upenn.edu\/~fdiebold\/papers\/paper112\/Diebold_Big_Data.pdf\" rel=\"nofollow\">originated<\/a> in the mid-90s, is a current buzzword amongst scientific communities. Rather than a sole reference to the size of complex datasets, the term broadly encompasses all aspects of working with large datasets from acquisition to analysis.<\/p>\n<p><b>Big data in science<\/b><\/p>\n<p>At its core, scientific research is driven by our curiosity to understand the relationship between cause and effect. Traditionally, \u2018hypothesis-driven\u2019 experiments are designed to answer a specific question about a cause-effect relationship.<\/p>\n<p>However, over the last sixty years there has been a <a href=\"https:\/\/gizmodo.com\/the-trillion-fold-increase-in-computing-power-visualiz-1706676799\" rel=\"nofollow\">trillion fold increase<\/a> in computing performance. The per-capita capacity to store information has <a href=\"https:\/\/www.sciencemag.org\/content\/332\/6025\/60\" rel=\"nofollow\">roughly doubled every forty months<\/a> since the 1980s. These technological advances are revolutionising almost all facets of human life, including how scientific research is conducted.<\/p>\n<p>In contrast to the traditional \u2018hypothesis-driven\u2019 approach, advancing technology allows us to acquire larger, more complex datasets, encompassing as many variables as possible, without bias from preconceived ideas. Powerful computation also enables us to finally realize the full potential of decades-old mathematical and statistical concepts. We can now sift through many variables and identify numerous cause-effect relationships in the same dataset, which would have previously been undetectable to the unaided human mind. These principles are now being applied to diverse fields, from astronomy to neuroscience, from particle physics to genomics.<\/p>\n<p><b>The need for collaboration<\/b><\/p>\n<p>The <a href=\"https:\/\/www.genome.gov\/sequencingcosts\/\" rel=\"nofollow\">National Human Genome Research Institute<\/a> reports that the cost of sequencing a human-sized genome was almost US$10 million in 2001, which had halved a couple of years later. The Human Genome Project took 13 years and cost about US$2.7 billion; however, human whole-genome sequencing is now more affordable and accessible than ever. Today, <i>Illumina<\/i>\u2019s <a href=\"https:\/\/www.illumina.com\/systems\/hiseq-x-sequencing-system\/system.html\" rel=\"nofollow\">HiSeq X Ten System<\/a> can sequence \u201cover 18,000 human genomes per year at the price of about $1000 per genome\u201d. Advances such as this have allowed scientists like Theordora Ross from <a href=\"https:\/\/www4.utsouthwestern.edu\/trosslab\/\" rel=\"nofollow\">UT Southwestern Medical Center<\/a> to identify novel mutations in \u201cmystery breast cancer patients\u201d \u2013 those with a strong family history of cancer but who did not possess the BRCA mutation \u2013 using human whole-genome sequencing. Advances in human whole-genome sequencing are also paving the way for <a href=\"https:\/\/www.nature.com\/ng\/focus\/icelanders\/index.html\">large-populations studies<\/a>, which in turn is inching us toward <a href=\"https:\/\/www.nih.gov\/precisionmedicine\/\" rel=\"nofollow\">precision medicine<\/a>.<\/p>\n<p>Thus, a lack of data is no longer the bottleneck to discovery. Rather, it is the effective management, analysis, and sharing of large datasets that now pose a challenge.<\/p>\n<p>Initiatives such as the <a href=\"https:\/\/www.opensciencedatacloud.org\/\" rel=\"nofollow\">Open Science Data Cloud<\/a> and the <a href=\"https:\/\/ns.umich.edu\/new\/releases\/23151-big-data-5m-to-widen-bottleneck-to-discovery\" rel=\"nofollow\">Multi-Institutional Open Storage Research InfraStructure<\/a> provide an online repository to efficiently store large datasets and share them between different groups. Effectively analysing complex datasets requires abilities that often extend beyond a single researcher\u2019s immediate skillset. Even the most tech-savvy researcher can struggle with some of the mathematical and computational expertise needed to correctly interpret large datasets. Thus, collaboration is key. Having a versatile team comprised of researchers, software engineers, bioinformaticians and statisticians, helps each focus on what they do best. There is no longer a requirement for the sole researcher to become a \u2018jack-of-all-trades\u2019. However, there is a need for clear communication between the experts of each field.<\/p>\n<p>Current global big data projects, such as the <a href=\"https:\/\/www.sdss.org\/\" rel=\"nofollow\">Sloan Digital Sky Survey<\/a>, the <a href=\"https:\/\/bluebrain.epfl.ch\/\" rel=\"nofollow\">Blue Brain Project<\/a>, and the Human Proteome Project, <a href=\"https:\/\/hapmap.ncbi.nlm.nih.gov\/\" rel=\"nofollow\">HapMap<\/a> effectively demonstrate the value of collaboration.<\/p>\n<p><b>Addressing the barrier to collaboration<\/b><\/p>\n<p>However, when it comes to projects that are being conducted on a smaller scale, many researchers are still apprehensive about openly sharing their data. The <a href=\"https:\/\/exchanges.wiley.com\/blog\/2014\/11\/03\/how-and-why-researchers-share-data-and-why-they-dont\/\" rel=\"nofollow\">reasons cited<\/a> include intellectual property concerns and the fear of being scooped. These concerns have been generated, in part, by the hypercompetitive environment of research, where a high impact factor publication alone has become the ultimate goal of scientists, no matter the cost.<\/p>\n<p>Journals such as <a href=\"https:\/\/www.nature.com\/sdata\/\"><i>Scientific Data<\/i><\/a> and <a href=\"https:\/\/www.gigasciencejournal.com\/\" rel=\"nofollow\"><i>GigaScience<\/i><\/a> help encourage researchers to share their data openly by recognizing their contributions as publications. Further, disseminating the entire dataset helps validate the interpretation of the data and the findings from it. It also opens the door to enable other researchers to reuse the data to investigate their own hypotheses, while guaranteeing proper acknowledgment of the source. For instance, different researchers can make maximal use of a large mass spectrometry dataset to investigate different proteins of interest, without the need for additional time and resources. This approach can help streamline scientific discovery with efficient use of funding.<\/p>\n<p>There are already discernible changes to the scientific research landscape that address the challenges of big data projects. However, the rise of data-intensive research requires a change of mind-set amongst scientists. There is an increased need for multidisciplinary research teams, with clear communication between experts of different fields. Scientists also need to be innovative and become more aware of the tools that will enable them to widely collaborate and openly share data. These changes will help us fully grasp the potential of big data and accelerate understanding.<\/p>\n<div id=\"attachment_7779\" style=\"width: 220px\" class=\"wp-caption alignright\"><a class=\"wpn-image-link\" href=\"https:\/\/blogs.nature.com\/naturejobs\/files\/2015\/09\/Lakshimi-mendis-naturejobs-blog.jpg\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-7779\" class=\" wp-image-7779 wpn-image \" title=\"Lakshimi-mendis-naturejobs-blog\" alt=\"Lakshimi-mendis-naturejobs-blog\" src=\"https:\/\/blogs.nature.com\/naturejobs\/files\/2015\/09\/Lakshimi-mendis-naturejobs-blog-300x300.jpg\" width=\"210\" height=\"210\" srcset=\"https:\/\/blogs.nature.com\/naturejobs\/files\/2015\/09\/Lakshimi-mendis-naturejobs-blog-300x300.jpg 300w, https:\/\/blogs.nature.com\/naturejobs\/files\/2015\/09\/Lakshimi-mendis-naturejobs-blog-150x150.jpg 150w, https:\/\/blogs.nature.com\/naturejobs\/files\/2015\/09\/Lakshimi-mendis-naturejobs-blog-1021x1024.jpg 1021w, https:\/\/blogs.nature.com\/naturejobs\/files\/2015\/09\/Lakshimi-mendis-naturejobs-blog.jpg 1077w\" sizes=\"auto, (max-width: 210px) 100vw, 210px\" \/><\/a><p id=\"caption-attachment-7779\" class=\"wp-caption-text\">{credit}Image credit: Lakshimi Mendis{\/credit}<\/p><\/div>\n<p><i>Lakshini Mendis is a winner of the <\/i><a href=\"https:\/\/blogs.nature.com\/naturejobs\/2015\/08\/12\/announcing-the-publishing-better-science-through-better-data-writing-competition\"><i>2015 Scientific Data writing competition<\/i><\/a><i>.\u00a0She is also a PhD student at the <\/i><a href=\"https:\/\/www.fmhs.auckland.ac.nz\/en\/faculty\/cbr.html\" rel=\"nofollow\"><i>Centre for Brain Research<\/i><\/a><i> in Auckland, and studies how the human brain changes in Alzheimer\u2019s disease. She is passionate about good science communication and is a strong advocate for women in STEM, and volunteers as the Editor-in-Chief at <\/i><a href=\"https:\/\/www.scientistafoundation.com\/lifestyle-blog\/introducing-stembox-a-monthly-science-box-for-girls\" rel=\"nofollow\"><i>The Scientista Foundation<\/i><\/a><i>! Follow her musings on <\/i><a href=\"https:\/\/twitter.com\/BLHSMendis\" rel=\"nofollow\"><i>Twitter<\/i><\/a><i>!<\/i><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Guest contributor Lakshini Mendis&nbsp; <a href=\"https:\/\/blogs.nature.com\/naturejobs\/2015\/10\/13\/big-data-collaborative-science#more-7777\" class=\"more-link\"> &hellip; Read more<\/a> <a href=\"https:\/\/blogs.nature.com\/naturejobs\/2015\/10\/13\/big-data-collaborative-science\/\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":45013,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1371,865],"tags":[429,17,993,1511,563],"class_list":["post-7777","post","type-post","status-publish","format-standard","hentry","category-competition-2","category-data","tag-big-data","tag-collaboration","tag-guest-contributor","tag-lakshimi-mendis","tag-scientific-data"],"_links":{"self":[{"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/posts\/7777","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/users\/45013"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/comments?post=7777"}],"version-history":[{"count":0,"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/posts\/7777\/revisions"}],"wp:attachment":[{"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/media?parent=7777"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/categories?post=7777"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.nature.com\/naturejobs\/wp-json\/wp\/v2\/tags?post=7777"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}