Showing posts with label Open Science. Show all posts
Showing posts with label Open Science. Show all posts

March 15, 2016

Preprint revolution in biology

A very nice article by Amy Harmon in the NY Times.

Handful of Biologists Went Rogue and Published Directly to Internet

On Feb. 29, Carol Greider of Johns Hopkins University became the third Nobel Prize laureate biologist in a month to do something long considered taboo among biomedical researchers: She posted a report of her recent discoveries to a publicly accessible website, bioRxiv, before submitting it to a scholarly journal to review for “official’’ publication.

...

And many #ASAPbio supporters retweeted John Hawks, a paleoanthropologist from the University of Wisconsin, who found himself recently at an African university where a paper on African genomes was unavailable because it could not pay the fee for the journal where it was published, and no preprint was available. He expressed his frustration with a profanity. 
... 
Preprint advocates counter that scientists care too much about their reputations to publish shoddy work, and posts to bioRxiv are clearly marked to indicate that they may contain information that “has not yet been accepted or endorsed in any way by the scientific or medical community.’’ Others note that plenty of peer-reviewed papers in high-profile journals have proved to be wrong, and some argue that carrying out peer review after a paper is published would provide a more rigorous and fair vetting of papers, anyway.  

July 22, 2013

New A00 project

A new project dedicated to Y-haplogroup A00 is seeking funding. A00 is the most basal clade of the human Y-chromosome phylogeny and has so far been found in African Americans and Mbo people from Cameroon. This is how the funds raised will be used:
In this campaign, we're seeking the funds we need to launch our first phase of fieldwork in late August, 2013. Our overall research plan includes five field trips to sample peoples in different regions of the country. In the first trip, Matthew will travel to the remote rural villages where he was raised, in the mountainous, forested Nkongho-Mbo region, and collect at least 100 samples from three villages, which will then be sent to a lab to be screened for A00. More in-depth testing will focus on those A00 samples. 
Planning and discussions are underway regarding which labs, probably more than one, will perform both SNP and STR testing on the samples. A full Y-chromosome sequence is on our wish list. This fundraising campaign, our first, will be limited to funding the DNA collection fieldwork. We'll be asking for your donations for lab testing in a separate campaign, in the near future. We need to limit the total amount of each fundraising campaign's goal to a modest amount we have some confidence of achieving within the set time, due to the all-or-nothing system used by Microryza.
 A00 is separated from the rest of mankind by >200 thousand years, but how closely related are different A00 chromosomes? It is necessary to have multiple samples to attempt to answer that question; I think it is quite probable that A00 is not limited to the area of Cameroon where it was discovered, but it probably makes sense to focus on it, since it is most likely to yield additional A00 samples given a sample size of 100.

A full Y-chromosome sequence of an A00 chromosome would also be very useful, since it would allow us to estimate its divergence from the A0-T remainder of mankind more securely.

The $2,500 they are seeking to launch their first field trip seems like a bargain to me.

May 08, 2013

The Geography of Recent Genetic Ancestry across Europe (Ralph and Coop 2013)

This paper first came out last July on the arXiv and went through four versions there before its final form which has now appeared in PLoS Biology. It's great that its early release allowed other people to read it without having to wait for the completion of the peer review process.

I think that this is a good model: journals have the right and obligation to subject papers to close scrutiny according to their own procedures, but this process ought not interfere with the early availability of research results or the ability of anyone other than the chosen reviewers to comment on new results.

PLoS Biol 11(5): e1001555. doi:10.1371/journal.pbio.1001555

The Geography of Recent Genetic Ancestry across Europe

Peter Ralph, Graham Coop

The recent genealogical history of human populations is a complex mosaic formed by individual migration, large-scale population movements, and other demographic events. Population genomics datasets can provide a window into this recent history, as rare traces of recent shared genetic ancestry are detectable due to long segments of shared genomic material. We make use of genomic data for 2,257 Europeans (in the Population Reference Sample [POPRES] dataset) to conduct one of the first surveys of recent genealogical ancestry over the past 3,000 years at a continental scale. We detected 1.9 million shared long genomic segments, and used the lengths of these to infer the distribution of shared ancestors across time and geography. We find that a pair of modern Europeans living in neighboring populations share around 2–12 genetic common ancestors from the last 1,500 years, and upwards of 100 genetic ancestors from the previous 1,000 years. These numbers drop off exponentially with geographic distance, but since these genetic ancestors are a tiny fraction of common genealogical ancestors, individuals from opposite ends of Europe are still expected to share millions of common genealogical ancestors over the last 1,000 years. There is also substantial regional variation in the number of shared genetic ancestors. For example, there are especially high numbers of common ancestors shared between many eastern populations that date roughly to the migration period (which includes the Slavic and Hunnic expansions into that region). Some of the lowest levels of common ancestry are seen in the Italian and Iberian peninsulas, which may indicate different effects of historical population expansions in these areas and/or more stably structured populations. Population genomic datasets have considerable power to uncover recent demographic history, and will allow a much fuller picture of the close genealogical kinship of individuals across the world.

Link

April 10, 2013

Closed-access story about DIY analysis tools

I find it a little odd that this story about DIY analysis tools, which (apparently) includes some quotes by myself, has now appeared in a closed-access publication. Had I known that to be the case, I doubt that I would have offered any response. It's probably not too late to make that item open access. 

In any case here's what I had to say (in full) to the author of the piece:

I think that a plurality of tools from a number of different analysts is an unambiguously good thing, both for the creators of these tools and their users.

For the users it is good because they can obtain different assessments of their ancestry, so they learn to be skeptical of extraordinary or unexpected claims of any particular test, and also to be more convinced of results that recur across many different tests.

For the creators it is good because of both (i) the motivation to improve their tools driven by competition with other test creators, and also (ii) the feedback they get from users of their tests.

These tools are also good for science in general, because a plurality of eyes (test creators and users) examine genetic data trying to detect interesting patterns in them that might be missed by more narrowly-focused research. So, a whole ecosystem of ideas springs up from these tests, as people try to fit their results into a broader pattern of human history. This is complementary to academic research: less structured and more "noisy" in terms of ideas that don't pan out, but also more dynamic, fast-paced and democratic.

As for Dodecad, I have developed my calculators by utilizing standard population genetics software, as well as software developed by myself, making use of publicly accessible academic datasets together with data from volunteers; the latter is very useful, because it helps me fill in gaps in population coverage: either because some populations have not been sampled in the literature yet, or, if they have, because their data is not publicly accessible to everyone.

November 16, 2012

TreeMix paper "officially" published

~8 months after the paper was pre-published in Nature Precedings, it is also "officially" published in PLoS Genetics. In the meantime, I count 18 uses of the label TreeMix in my blog, which includes both uses of the treemix software itself and its auxiliary threepop and fourpop programs; I also wrote a small script that converts ADMIXTURE output into TreeMix format, and generally had a lot of fun using it. I'm glad I didn't have to wait 8 months to learn that something like TreeMix existed.

In the grand scheme of things, an 8-month head start may not be much, but consider that perhaps someone else might either have a use for TreeMix or the desire to build on it, and if they decide to make their research available prior to official publication, then, perhaps an additional few months might be gained. And, if someone else still decides to follow up on them then...

There are many arguments for immediate publication of research results, but I think that the potential for speeding up scientific progress is one of the best ones.

In the old days, it was really necessary to impose a delay between the time when a scientist placed a final full stop to his paper and the time it appeared on another scientist's desk: publication involved significant expenses of paper, ink, and labor, so the frivolous or erroneous had to be weeded out; dissemination involved expensive transport by carriage or boat; storage involved a building, and bookshelves, and additional cost.

All these costs have shrunk to insignificance; imposing delays to research dissemination now accounts to little more than placing a sleep() call in the unending loop of scientific advancement. And, the one remaining argument for post-review publication ("weeding out the frivolous or erroneous") carries little weight: pre-review publication is a better guarantor of quality by exposing research to many more eyes and minds that may scrutinize it more carefully, having rid themselves of the idea that "if it's published it must be good".

PLoS Genet 8(11): e1002967. doi:10.1371/journal.pgen.1002967

Inference of Population Splits and Mixtures from Genome-Wide Allele Frequency Data

Joseph K. Pickrell1, Jonathan K. Pritchard

Many aspects of the historical relationships between populations in a species are reflected in genetic data. Inferring these relationships from genetic data, however, remains a challenging task. In this paper, we present a statistical model for inferring the patterns of population splits and mixtures in multiple populations. In our model, the sampled populations in a species are related to their common ancestor through a graph of ancestral populations. Using genome-wide allele frequency data and a Gaussian approximation to genetic drift, we infer the structure of this graph. We applied this method to a set of 55 human populations and a set of 82 dog breeds and wild canids. In both species, we show that a simple bifurcating tree does not fully describe the data; in contrast, we infer many migration events. While some of the migration events that we find have been detected previously, many have not. For example, in the human data, we infer that Cambodians trace approximately 16% of their ancestry to a population ancestral to other extant East Asian populations. In the dog data, we infer that both the boxer and basenji trace a considerable fraction of their ancestry (9% and 25%, respectively) to wolves subsequent to domestication and that East Asian toy breeds (the Shih Tzu and the Pekingese) result from admixture between modern toy breeds and “ancient” Asian breeds. Software implementing the model described here, called TreeMix, is available at http://treemix.googlecode.com.

Link

October 12, 2012

From Skulls and Scans (Monge & Aguirre @ Penn)

An interesting talk on the uses and abuses of science.

 

I find this excerpt particularly offensive to my open science sensibilities (29:40 ff):
We were able to, after 10 years, actually get it published in PLoS Biology, took us 10 years, yeah, we were rejected every single place that we ever sent the manuscript. I find that interesting too, and we even had editors say to us that they didn't want to say anything against Stephen J Gould. Isn't that interesting?
Interesting indeed. As I wrote in my review of the Lewis et al. (2011) paper on the Gould vs. Morton affair:
It is remarkable that 30 years after the Mismeasurement of Man Gould's errors are uncovered. Why did it take so long? While one could understand why the (totally unfounded but -on the surface- plausible) idea of measurement bias could have gone unnoticed until someone actually re-measured the skulls, but the statistical error that Gould committed was there for anyone to see.
We now know why it took at least 10 out of these 30 years: stuck in journal limbo.

August 19, 2012

Raising a peace banner in the Neandertal Wars

The two camps in the Second Neandertal Wars (*)  have assumed maximalist positions on opposing sides of the argument: African structure explains it! vs. Neandertal admixture explains it!. Armed with the Vindija genome, that marvel of technological ingenuity, and a suite of impressive statistical models, the two sides have reached completely opposing conclusions.

In order to formulate my own position, I decided to do what I love best, i.e., to look at the data for myself. My main idea is that the signals of Neandertal and Denisova admixture as measured by these quantities (D-statistics) ...

D(Pop1, Yoruba, Neandertal, Chimp)
D(Pop1, Yoruba, Denisova, Chimp)

... will vary on different SNP ascertainment panels. SNPs ascertained in Africans may have a great number of Palaeoafrican alleles; SNPs in Neandertal-admixed populations will have a great number of Neandertal alleles; SNPs in Denisova-admixed populations will have a great number of Denisova alleles. If a population has admixture from hominin X, this admixture, as measured by the D-statistic, will tend to be inflated in panels possessing alleles that introgressed from X, and suppressed in panels that lack them.

The issue of ascertainment and archaic admixture was addressed by Skoglund and Jakobsson (2011); my aim is different: I am not so much interested in how ascertainment affects admixture estimates, but rather in exploiting the observation of the preceding paragraph (that Palaeoafrican, Neandertal, or Denisovan SNPs will lurk at different rates when ascertained in different individuals) to see what it tells us about human differences.

The signal of "archaic admixture" may be generated by genuine archaic admixture in one population (e.g., Eurasians), making it more similar to the archaic group (e.g., Neandertals), or by archaic admixture -of a different sort- in another population (e.g., Africans), making it less similar to that group. Both these processes may be at work, operating at different intensity in different populations and across different timelines.

I used the Harvard HGDP set, which contains 12 SNP panels, each of which has been ascertained in two chromosomes of a single individual. These panels are:
San, Yoruba, Mbuti, French, Sardinian, Han, Cambodian, Mongolian, Karitiana, Papuan1, Papuan2, Melanesian
A D-statistic was calculated relative to either Neandertal or Denisova for all HGDP populations, as well as the two archaic hominins. Subsequently, I used MCLUST to infer the number of different clusters on the basis of these statistics. In the optimal solution, MCLUST inferred 7 clusters, with each archaic hominin getting its own cluster, while the modern human populations were assigned to 5 clusters corresponding to five major human races recognized by traditional physical anthropology (Mongoloid, Negroid, Australoid, Capoid, and Caucasoid).
Note that these are not admixture proportions, but assignment probabilities! All populations fell into their expected clusters. The populations from Pakistan who are believed to be predominantly Caucasoid with varying degrees of minor admixture of an Ancestral South Indian element were assigned to the Caucasoid cluster. So did the Mozabite Berbers, a Caucasoid population with minority Negroid admixture. Finally, of the Central Asian populations, the Hazara of Pakistan showed mixed affiliations in the Caucasoid and Mongoloid clusters, while the Uygur were assigned to the Mongoloid cluster.

It is noteworthy that by exploiting patterns of relationship of modern to regional archaic humans, we have managed to recreate the major human groups. This is, perhaps, supportive of those who have argued that a degree of regional continuity across the Old World, and not only recent post-Out of Africa genetic divergence is responsible for present-day inter-population differences.

MCLUST also gave us the D-statistic means for the 7 inferred clusters. Remember that these are differences between a population Pop1 and Yoruba, relative to an archaic hominin (Neandertal or Denisova), and for 12 different ascertainment panels:


There are wonderful patterns to be discovered here; you can look at the data for yourselves; that's the open science thing to do.

All our ideas about human origins are conditioned on the availability of genomes from two archaic Eurasian hominins, and the lack of genomes of similar age from Africa.

But, remember:
  • You can fit Europe, China, India, and the US into Africa, with room to spare. 
  • If Vindija and Denisova, two caves less than 5,000km apart were home to people more divergent from each other than any two humans are today, it's strange to think that only "modern humans" inhabited Africa at the same time. 
  • The maximum genetic distance between living Africans is much higher than the maximum distance between living Eurasians: Africa is much more diverse than Eurasia. It's simpler to assume that the same relative pattern was true during the Middle Stone Age. The palaeoanthropology seems to support this, showing archaic forms present even during the terminal Pleistocene in Africa.
  • If modern humans did interbreed with 2/2 archaic humans whose sequences we possess, it's strange to think that they somehow shunned the African Others.
In view of the above, I humbly raise my peace banner in the Neandertal Wars, and declare that it isn't either-or: it's both!

(*) The First Neandertal Wars were fought decades ago by anthropologists working with calipers and magnifying lenses. Their outcome was to relegate Neandertals from the enviable position of our likely ancestors to that of an irrelevant sidekick, although a not-negligible minority continued an insurgency against the Out-of-Africa-only victors.

Four steps to open science


As an advocate of open science, I think there are four steps in the process toward a really open scientific culture:

  • print journals were the first, early, step, because they allowed dissemination of scientific knowledge to a wider audience. They were the killer app of the 17th century, combining the invention of the printing press with the idea of periodically bundling up new ideas and disseminating them to anyone who cared in one package.
  • the "open access" movement is a second step, because it removes the monetary barrier to knowledge acquisition. Modern journals need not be printed or paper or transmitted on wheels: they can live as bits in abundant persistent storage media, and be transmitted over high speed communication lines. The cost of assembling and disseminating a new idea is negligible.
  • the "pre-publication review" movement is a third step, because it de-privileges a limited set of reviewers and makes research results available earlier. The chain of scientific progress can have shorter links (because people are aware of- and build on new ideas earlier), and more eyes and brains can scrutinize new ideas. Journal editors of old had to seek expert of opinions of a few, but thanks to the wonderful invention of costless one-to-many broadcasting, there is no reason to rely on the few, rather than the many.
  • the replacement of the "article unit model" with an "open-source" science model in which knowledge is assembled from bits and pieces from thousands of sources, constantly updated, constantly reviewed, constantly tested against new evidence.

The last step may be the most difficult, because it goes so much against the culture of competition between individual scientists and research groups. Why do people wait on new ideas and new results? Because they either want to assemble enough material for an LPU, or develop their ideas exhaustively to merit publication in a prestige journal.

What if Newton had published a couple of paragraphs on his calculus idea in 1666 and not decades later? What if he were able to tweet his apple incident and not wait to publish his Principia? What if Darwin had not delayed his publication of Origin until he was afraid of being scooped, but had mentioned the idea decades earlier when he conceived them? Perhaps, someone else would have taken their ideas and run with them, and we'd be living in a 2050s level of technological progress today. Perhaps, if others who followed them did the same, technology on earth would have advanced by centuries relative to its current level.

There are limits to our ability to digest and build on new ideas, but I would argue that getting rid of the "article unit model" and adopting an "open-source" attitude would accelerate the pace of scientific progress. Journal articles won't disappear, but rather than being at the vanguard of progress, they will be at its rear, like stable releases of open source projects that weed out the bad, keep the good, and give everybody a reference point to work against.

August 01, 2012

Let's play ASHG 2012 title imputation! (+open science miscellanea)

It's that time of year again, and the titles for the ASHG 2012 presentations have just been posted online. Well, part of them anyway; the dreaded (...) have made another appearance. Still it is fun to try to guess what each contribution is about, and we'll only have to wait ~1 month for the abstract text.

In related news, Ewen Callaway reports on the trend (?) for biologists to put their unpublished work in arXiv. My own views are strictly for open science, so I applaud the people who are dragging their disciplines into the 21st century.

Finally, recent initiatives in the UK and the EU will mandate open access for work funded by research agencies. This is a good step in the right direction, but a very incomplete one: open access solves the problem of ensuring wide dissemination of new science, but merely shifts the flow of public  money rather than sever it. With open access Government->University Library->Journal is replaced by Government->Research Agency->Scholar->Journal.

Moreover, open access does not address the more fundamental issue of how journals impede scientific progress by imposing the antiquated pre-publication peer review process. The sky hasn't fallen over the heads of physicists who post their work on arXiv when they're done with it and carry out post-arXiv publication peer review. So, it probably won't fall on the heads of biologists who do the same either.

A good example of this is the recent work on ChromoPainter/fineSTRUCTURE that appeared months before publication: lots of people -including myself- started using their software right away, which spurred new insight, and they got their peer-reviewed publication too. More recently, a group of independent researchers co-ordinated their efforts in public to hack 1000 Genomes data, discovered and validated new SNPs, and they got their publication too. Open science works, so everyone should try it!

June 06, 2012

How journals once facilitated and now hinder scientific progress.

Scientific journals, were instrumental in the ascent of a scientific culture during the modern era. Before their invention in the 17th century, there was, of course, communication of novel scientific findings. However, this largely took the form of:

  1. visits between scholars, not particularly easy by steam and coach, 
  2. letters between scientists interested in the same topics, 
  3. ad hoc collections of writings in the form of manuals, lecture notes, or 
  4. more organized publication of monographs for mature and complete work.

Science during the pre-journal era was not widely disseminated. It was the privilege of the few who aggregated in university towns or around noble patrons. There was, indeed, a culture of distrust, the most famous episode of which was the infamous feud between Newton and Leibniz on the invention of calculus. And, indeed, even though there was already an emergent bourgeoisie with scientific interests, these were largely excluded from science by the mere fact that they were "disconnected" from the network of scientists.

Scientific journals solved many of these problems:

  1. They introduced a centralized method for assigning credit to authors (via publication) and disseminating science to all interested parties (libraries, wealthy dilettantes, etc.)
  2. They made science communication and progress significantly faster, since novel findings could be revealed in piecemeal fashion, and widely disseminated to anyone who might be interested in them and/or build upon them.


Scientific journals made the dissemination of science faster, cheaper, and wider.

I will argue that modern journals perform the exact opposite function: they cause the dissemination of science to be slower, more expensive, and more limited.

How can I argue that modern journals make the dissemination of science slower? The proper comparison is not, of course, against the speed of dissemination of the 17th century, but, rather what is technically possible at any given time.

When journals were invented, they more or less allowed results to be made available as soon as possible, given the limitations of wind, horse, paper, ink, and the necessity of combining contributions into collections that could be printed periodically.

Contrast with the present day: scientific contributions could become available to billions of people, as soon as they have been written. The only real remaining obstacle is the brain-hand-computer interface. But, this, is not, of course, what takes place: rather:

  • scientists postpone the publication of preliminary results until they have a cumulative piece of work that is "publishable" according to journals' standards
  • they submit their work to a closed system of peer review in which a handful of eyes decide whether their work merits publication or not. This takes time, and withholds the work from judgment from literally everybody who might have something to say about it, professional scientists and laymen alike
  • they journal-shop their contributions, with perhaps several rounds of submission/review/rejection/re-writing/re-submission. This may, of course, make their work better, but it introduces months if not years of latency; papers could, in fact, be improved and enhanced post-publication.
  • finally, they must abide by journals' rules regarding publication. The necessity for bundling up contributions in paper format has largely disappeared, but editors' need to plan issues and schedule papers and "publicity" accordingly has not.

Let's proceed to point #2: journals make the publication of science more expensive.

Again, the proper comparison is not with the 17th century, but rather with what is technically feasible today: we can now disseminate scientific research entirely for free.

If you've read this blog for a few years, neither I nor you have paid a cent for it. It could be argued that someone (i.e., Google or other service providers) is paying for this ability, but they are not forcing anyone to contribute. Obviously, they have a business model which allows them to provide free bandwidth to millions of content generators. No one has to pay to read/publish text on the web.


Transmitting papers over the web has a trivial cost. This cost is, indeed, so small, that it is probably negligible in the context of running the Internet. By their very nature, papers are usually small compared to multimedia files. And, even if they became much larger and included lots of multimedia supplements, they could still be easily transmitted using peer-to-peer file sharing technologies and the like.

On the contrary, journals today charge for access to their papers. Sometimes, it is subscribers that pay, else it is authors (in many "open access" journals). But, what is the service that journals actually provide to science? Authors can easily format manuscripts to a readable template, and reviewers contribute their time for free. Journals are leeches on society that provide no discernible service.


At most, it could be argued, that the hierarchy of journals allows everyone to assess the value of new scientific findings. If it's published in XYZ, the idea goes, it must be important. But, this is a pernicious idea for at least two reasons: first of all, there is no shortcut for assessing the value of a scientific paper: you have to read it!

Second, the air of respectability that journal publication provides allows bad science to survive and replicate. Once in a while, some story comes along, such as the "arsenic life" debacle to remind us that being published in XYZ isn't a guarantee of merit. And, indeed, what better way to facilitate the "weeding out" of bad science than to allow more eyes to scrutinize it?


Science dissemination is now more limited. In this respect, the "open access" model has a supposed advantage over the "closed access" one, since it allows everyone to read a paper.

But, the leading variety of "open access" -in which authors pay for publication- is also limiting. It does so not by limiting who gets to read what is published, but by limiting who gets to publish. The choice cannot be for authors to either "pay up" or for readers to "pay up". No one really has to pay anything more than is necessary to run and maintain the Internet infrastructure itself.


Conclusion


I hope I have convinced you that modern science publishing is utterly indefensible: it delays and limits the dissemination of science. It prohibits the poor from publishing or reading science, and hence it helps reinforce poverty and ignorance.

Who benefits from the current model? Of course, rent-seeking journals do. But, we should not absolve scientists themselves of their responsibility. The journal system is scientists' way of assigning and receiving credit for their work. It is a sine qua non of the modern professional science world with its emphases on careerism, grant-seeking, etc.

We must not be oblivious to the fact that what may be good for "scientists" may not be good for Science, and has indeed become a poison that is threatening to strangle Science itself. Can there be redemption for this culpability? Scientists, after all, are supposed to be innovators, so it is surprising how reactionary they have become to any changes in their modus operandi. 

More and more people seem to understand the problem, but few are willing to embrace the solution. It may be ironic, but the road to scientific freedom may ultimately depend on government coercion precipitated by democratic demand. I don't hope that this will happen soon, both because of corrupt governments and an apathetic populace, but it must.

21.9% of human variation datasets withheld

PLoS ONE 7(6): e37552. doi:10.1371/journal.pone.0037552

Mine, Yours, Ours? Sharing Data on Human Genetic Variation

Nicola Milia et al.

The achievement of a robust, effective and responsible form of data sharing is currently regarded as a priority for biological and bio-medical research. Empirical evaluations of data sharing may be regarded as an indispensable first step in the identification of critical aspects and the development of strategies aimed at increasing availability of research data for the scientific community as a whole. Research concerning human genetic variation represents a potential forerunner in the establishment of widespread sharing of primary datasets. However, no specific analysis has been conducted to date in order to ascertain whether the sharing of primary datasets is common-practice in this research field. To this aim, we analyzed a total of 543 mitochondrial and Y chromosomal datasets reported in 508 papers indexed in the Pubmed database from 2008 to 2011. A substantial portion of datasets (21.9%) was found to have been withheld, while neither strong editorial policies nor high impact factor proved to be effective in increasing the sharing rate beyond the current figure of 80.5%. Disaggregating datasets for research fields, we could observe a substantially lower sharing in medical than evolutionary and forensic genetics, more evident for whole mtDNA sequences (15.0% vs 99.6%). The low rate of positive responses to e-mail requests sent to corresponding authors of withheld datasets (28.6%) suggests that sharing should be regarded as a prerequisite for final paper acceptance, while making authors deposit their results in open online databases which provide data quality control seems to provide the best-practice standard. Finally, we estimated that 29.8% to 32.9% of total resources are used to generate withheld datasets, implying that an important portion of research funding does not produce shared knowledge. By making the scientific community and the public aware of this important aspect, we may help popularize a more effective culture of data sharing.

Link