Effectively eDNA observations - what to do?

The recent large volume of North American fungal observations identified by sequence barcoding has created some interesting challenges and here is one I think needs some discussion. Observations are appearing of mushrooms not identified as the photographed mushroom but as a contaminant that appeared in the sequence data. The barcode sequence is pasted into an observation field. That contaminant (often obscure microscopic yeast species) is not visible in the photos, but will almost certainly be there. DNA extracts from mushrooms taken from the environment will almost always contain dozens of species, but the methods used for sequencing and analyzing the data usually pick out the intended target and ignore the rest. To my mind these observations identified as ‘the rest’ are effectively eDNA samples. For example how would it differ from uploading the same photo of a teaspoon of soil creating hundreds of identical observations with the set of with eDNA-identified species in the same sample? I don’t think that would be acceptable for iNat.
My question is whether the DQA flag “doesn’t present evidence of an organism,” should be used? In one sense the evidence is there because of the pasted sequence. But, if they are allowed then they have the potential to become RG, and associated with photos that in no way resemble the species, thus confusing both end users and the CV. On the other hand there are plenty of RG observations of fungal disease symptoms that also don’t show the causal fungal agent and where often symptom characteristics are not uniquely diagnostic (but that’s different issue).

As an identifier, how can we realistically confirm or deny DNA? Would confirming require you yourself doing DNA extraction and the whole process?

Also I’m not sure if i hate or am neutral about all these new fungi IDed solely from DNA getting in the CV. I think if it continues fungi species suggestions in the CV could go down a path to becoming near useless in many cases, although i can imagine genus or higher taxa suggestions actually getting more accurate with the inclusion of the new DNA fungi.

Eitherway I’m not confident the CV was built with tons of DNA IDed provisional fungi. Also why do provisional fungi get to be learned, but not hybrids?

Also from what I’ve seen in general. Identified by DNA tends to illicit many blind agreements.

Photos or sounds attached to observations should include evidence of the actual organism

https://inaturalist.freshdesk.com/en/support/solutions/articles/151000169928-what-kind-of-photos-or-sounds-should-i-attach-to-observations-

This doesn’t seem to be met by observations of environmental microorganisms like you mention. It’s essentially including only a habitat photo that doesn’t include the organism, as the photo would not look any different without the organism.

It’s a little unclear whether this means there is “evidence” of the organism or not, i.e. do observation fields or other text count as evidence when the photos don’t. I would say no, as an observation without photos is casual despite a very detailed description of the organism. But that’s technically not explicitly stated in the FAQ.

I don’t want to conflate two separate issues. My question was specifically about DNA based identifications (of formally described taxa) without supporting photos showing the target organism.

The application of (and problems with) provisional names is a different issue. In general identifications of formally named species supported by barcodes (and the relative objectivity that provides) should significantly improve the accuracy of the CV for fungi. It certainly needs it. In general too few fungal names are being applied to too many obviously different species (in the northern hemisphere) leading to poor CV suggestions.

I really like “as the photo would not look any different without the organism”. That nails the problem for me.

From my perspective with the way iNat is designed (the assumptions of the observation UI, how the image galleries and CV are programmed, etc.), the media uploaded to an observation (photos or audio) are the primary evidence and anything in the observations notes, comments, or observation fields are supplementary evidence. Consequently, when in doubt, I would use the information in the primary evidence to make my ID.

For example if I were identifying an observation of a bird and it had a photo of one species and a note perfectly describing the song of another species, I would identify the species in the photo and add a comment indicating what other species the note describes. In the case in question if there’s one fungus shown in the photo and a different one in the observation field then ID the one in the photos.

This was my argument here which I think was misunderstood. I trust many observers to provide reliable field notes, but in the context of iNat those aren’t considered sufficient for ID confirmation without media evidence. To be consistent with the assumptions of iNat’s system overall, supplementary evidence (as I defined it above) is useful for increasing confidence in my identification but not sufficient by itself to make an identification without support within the primary evidence. I suppose DNA evidence is theoretically a uniquely objective form of evidence so I get going further with IDs from that than just from observer descriptions.

I think this was better expressed in this post, and Reuven’s wording around the “specimen in question” implies that the photographed mushroom is assumed to be the intended organism for identification:

I think this is also consistent with our general strategy for identifying Unknowns, where we choose whatever we want unless the observer indicates what organism they intended the observation for. If there’s a photo of a mushroom and not a microscopic yeast then presumably the observer is interested in the mushroom.

I agree that, if there is a conflict between the fundamental observation data (i.e., media) and data in a supplemental/ancillary field, then the core media takes precedence. So I would ID for the organism in the media.

I think @thomaseverest 's point is a good one - iNat doesn’t accept text-based descriptions of organisms (i.e., those without media) as eligible for research grade. A DNA sequence on it’s own isn’t eligible for RG, and that is the only specific evidence of the specific organism sequenced being provided in an observation like this.

That’s not to say that the DNA sequences aren’t valuable data - they just aren’t a good fit for iNat. They should go to Genbank, etc.

I would also note that, without other evidence, it’s impossible to verify that the DNA sequence is actually present even in the “habitat” photo. It could easily have come from equipment used to collect the sample, non-sterile technique, lab contamination, etc. This happens reasonably frequently with DNA samples collected under field conditions. Just a DNA sequence of a microscopic organism that can’t otherwise be attributed to a specific location isn’t strong evidence it was actually at the site of the observation and that the DNA wasn’t introduced at some other point in the process.

Jerry, I know you said:

But the problem you describe will become more common, and to speculate, we might all eventually have DNA samplers on our phones, so I will add that if the sequence could be combined with microscopy, especially marked up with circles and arrows showing the differences between the different types of cells, this would be visual “evidence of organism”.

It requires more work from the observer, but I guess it’s justified in this use case.

It is common practice with fungi (macro or micro) to provide microscope photos, with or without sequence data.

Mushroom microscopy · iNaturalist

One problem is that iNat provides no ability to annotate these photos individually, so there is no direct way of annotating these micro photos, even just to tag them as microscope photos so the CV training doesn’t try to ingest them. The lack of ability to annotate individual photos is a serious deficiency with iNat for microscopy. It was a request I made perhaps only a few years after iNat started, but it was declined.

If the photo shows the organisms and has a DNA squence, I’d say accept it, even if the mushroom itself isn’t identifiable. Not necessarily ID it, but certainly don’t score to disapear. There is evidence of the organism in question.

I think the situation is different if the DNA is from some parasite or commensal what lives in the organism shown and can’t be seen in the photo. Then, I agree, it’s like a habitat photo without the organism showing. So my initial response is mark it “No evidence of organism.”

What about the fact that the DNA sequence is there? That makes it harder. After all, somebody who knows how to read this DNA sequence can confirm of deny that it represents the invisible organism. But we’re not seeing or hearing that organism. Hmm. This doesn’t fit the iNaturalist situation right now but ought to be discussed. I think iNaturalist needs to make a decision about this and write it up.

Why is the potential that the CV might learn microscopy photos a problem? If microscopy is an important way to identify a species, wouldn’t you want it to learn how to distinguish microscopy photos of different species?

For statements like these two:

I’d like to refer to this excellent paragraph here:

As identifiers who are given just a photo and a DNA barcode we have no way to independently verify the link between the two. If the uploader mixes up their samples/photos and mistakenly uploads completely unrelated visual and DNA data there is no way to tell.

Another thing we need to consider is that the person who made the observation and collected the sample is not always the one doing the testing or making the DNA based ID. If they observe a visible mushroom and send it in to see which cryptic species it is, and the results come back for a microscopic yeast, the observation is still of the mushroom they photographed. I would identify the mushroom.

If observer and tester are one and the same, and they want to observe the yeast, then they should culture it out and take microscope photos. Without that, maybe respect their intention that it’s for the yeast, but mark no evidence of organism? (I haven’t actually encountered this, and might just mark reviewed and move on if I ever do.)

I think in those cases the symptoms are visible "evidence of" the organism that caused them, just like we can identify a deer track as a deer, or a gall as the wasp that made the plant form the gall. Sometimes the combination of disease symptoms and host plant are enough for a species level ID of “what made these spots on this leaf.”

My statement there was in the context of my general position on supplemental evidence. In other threads on this forum I have been more skeptical than most people in trusting observers’ notes to confirm a species when the provided photos are insufficient. For example if there is a photo of a crow or a flycatcher that is only identifiable to genus from the photos, and the observer adds “I heard the species’ distinctive call/song” I would likely still only identify it to genus because I don’t know if I can trust the observer to reliably identify the call/song, and they didn’t upload audio to allow me to independently confirm it.

However if I were to extend this logic to DNA evidence then it would defeat the purpose of the DNA analysis that many fungi observers and identifiers rely on. “Maybe there’s a chance that there are two species of mushrooms in the same genus right beside each other and the photo is of one and the DNA could be of the other, so I can only identify to genus.” That seems plausible but not likely enough to be concerned about. A similar example is if a leafminer or gall identifier relies on knowledge of the plant host to identify an observation, and the photos of the leafmine/gall don’t show sufficient details of the plants to identify it (although ideally additional photos would be included in a separate observation of the plant in this case).

Because I think you would deliberately need to partition the training photos into sets of photos showing the same characters, not mix them all together.

At face value, you are correct, it shouldn’t be an issue and seems positive. However, what good is training data if 99.9% of users do not upload microscope images? For my group, I can probably count on both hands how many people upload male genitalia under a microscope, positioned correctly, and not shriveled up and distorted.

The CV is already trained on taxa that have multiple different forms – e.g., sexual dimorphism; radically different juvenile and adult forms; tracks, nests, feathers and other non-organism forms of evidence, etc.

It is an image-matching algorithm. It doesn’t know what it is seeing and this doesn’t necessarily matter.

What is the harm?

Single images will likely be treated as outliers in the learning process – if they end up in the selection of photos used in the training at all because they make up such a small percentage of the potential training material. This is not an argument for actively excluding such images from eligibility for CV training. It is, if anything, an argument for encouraging users to include microscope images if they are able to do so.

I agree. I don’t know if this is too off-topic for this particular thread, but the influx of fungus DNA-based IDs over the last few years is giving me flashbacks to the Lepidopterist world of the early 2010s. For anyone reading who isn’t aware, DNA “barcoding” involves sequencing a small piece of DNA that’s been found to generally vary a lot between species and not much within species. For Leps (and I think for all animals), we use roughly 650 base pairs from cytochrome c oxidase I mitochondrial DNA, and my understanding is the for fungi they’re using a similar length of nuclear DNA from a spacer region of the DNA that codes for building ribosomes called ITS (internal transcribed spacer). While COI and ITS barcodes are very useful, there are some massive limitations to using just this barcode, in the absence of any other evidence, to identify a specimen or change taxonomy:

  • It’s not a genome. It’s a very tiny sequence of DNA from one specific place. It’s a useful bit, don’t get me wrong. There’s a reason these are what we use for barcoding, and there are papers out there estimating that well over 90% of species in certain clades can be differentiated by their barcodes. But that still leaves many thousands of species that can’t.
  • Barcodes are not identifications. BOLD matches unknown sequences to sequences that have been assigned a name. If the BIN (barcode identification number) that matches your sequence has been assigned a name, then the obvious question is “who gave it that name?” For Leps, there are countless BINs with incorrect names on them, because the first specimen of that BIN sequenced was misidentified. So people sequencing specimens with that BIN now confidently declare that their specimen was “confirmed by DNA sequencing” to be species X, which is in fact just repeating someone else’s misidentification from several years ago.
  • DNA barcode is one piece of evidence to consider when splitting/lumping species, but it’s not the end-all. There are many examples of sister species with identical barcodes that are clearly acting as different species by any reasonable species concept. There are examples of species with a surprisingly high level of barcode variation. Peer reviewed publications that consider barcode alongside ecology, biogeography, anatomy, phenology, etc. can leverage the barcode as one piece of the puzzle. Someone armed with a couple records of specimens with a novel barcode sequence should, IMHO, not be given a platform to declare that they’ve found a likely new species.

What happened in the Lepidopterist world was that thousands of photos got uploaded onto BugGuide and Moth Photographers Group with IDs that were “DNA confirmed”. Species pages were added for unique BINs that contained “probable new species” based on the newest barcode results. There were people pushing the lumping species if their barcodes were indistinguishable. In short, hobbyists got way too excited about DNA sequencing and took things too far.

To see what the Lep world looks like in the aftermath of all this hype, look at Moth Photographers Group, where all the species pages still have a link to BOLD, but with the warning label “Caution: Identifications often erroneous; DNA barcode provides evidence of relatedness, not proof of identification”. The exponentially-increasing “species pages based on unidentified BIN” were halted in favor of a more conservative approach that waits for peer review before changing things. Barcoding is still an integral part of Entomology, and I can’t imagine doing modern taxonomy without it. But it’s up for debate whether the wild west hobbyist barcoding phase we went through resulted in more clarity or more confusion. I spent the better part of a decade getting chastised for daring to correct misidentifications on BugGuide by users who were certain of their ID because “barcode”. We’re still finding obvious misidentifications that need fixed from 2009-2019 where the reason for the original ID is given as “DNA” with a link to BOLD.

I know the narrative is wildly appealing: “Putting the power of DNA testing into the hands of the people so we can work out taxonomy and confusing species on our own without waiting for the slow progress of the old guard of stuffy scientists in their ivory towers!” I get it. It seems like anyone pushing back against such a trend must be an old elitist curmudgeon who wants to hold onto their power and influence over the community. But having gone through the past 20 years as a Lepidoptera hobbyist, I worry that the fungus folk will make some of the same choices on iNat now that we did on BugGuide 15 years ago, and it’s concerning. We’ve already started a pilot “add your own species ahead of the published taxonomy based largely on barcode” program, and after reading the long thread on that announcement, it’s clear that it’s immensely popular among mycology enthusiasts. I did find it telling though that on that thread, the techniques were framed as something that follows in the footsteps of what Lepidopterists have been doing for many years. Yet it’s also pointed out that the DNA barcode sequence field on iNat is used almost exclusively by mycology folks. Why aren’t the Lepidopterists pushing so hard for barcode-based IDs and species pages? Aren’t we the ones that started this trend? Why did we pump the breaks on it? …

I think I know what to do about “DNA” Identified observations in my group on iNaturalist instead of just ignoring them since they practically never have photos showing the morphological features needed. I’m not actually sure if there is any DNA Identified chiro observations on the site that also has the accompanying appropriate morphology. Nothing comes to mind off the top of my head.

Though I think another question to be raised. Is it blindly agreeing if a user sees DNA, and agrees? This then would make it a moderatable thing if a user just goes through clicking agree because “DNA”, thereby making those observations RG.

Again, I’m not sure anybody has explained how do you as an identifier confidently and independently verify an observation based on DNA.

If somebody sequences a chiro, uploads one image of it not showing the relevent morphology, I can’t confirm it.

If you go to https://id.boldsystems.org/ you can enter a sequence and compare it to any of the public barcode databases. For example, here’s a barcode sequence for a random Chiro:

TCTTTACATTATTTTCGGTGCTTGATCAGGAATGGTAGGGACTTCTCTAAGAATGCTTATTCGAGCAGAATTGGGTCGACCTGGGACTTTCATTGGTGATGACCAAATTTATAATGTAGTAGTTACAGCCCACGCATTTATTATAATTTTTTTCATAGTAATACCAATTTTAATTGGTGGTTTCGGTAATTGACTTGTACCTTTAATACTTGGGGCCCCAGACATGGCTTTCCCCCGAATAAATAATATAAGTTTTTGACTTCTTCCCCCTTCTCTCACTCTTTTACTTTCTAGTTCATTTGTAGAAAATGGAGCAGGAACAGGTTGAACTGTTTATCCCCCTCTTTCAGCAGCAATCGCCCATAGTGGAGCTTCTGTTGATTTAGCTATTTTTTCTTTACATTTAGCGGGTGTTTCTTCTATTTTAGGATCTGTAAATTTTATTACTACAGTAATTAATATACGAGCAAATGGAATTACTCTAGACCGAATACCTTTATTTGTTTGATCAGTTGTCATTACAACTGTTTTACTTCTTCTTTCTTTACCAGTTTTAGCAGGTGCGATTACTATACTATTAACAGACCGAAATTTAAATACATCATTTTTTGACCCAGCCGGGGGTGGTGACCCAATTTTATACCAACATTTAT

I selected “ANIMAL LIBRARY (PUBLIC)” and pasted the sequence in the space below the databases. It did a search, and found 4 identical barcodes, 3 of which were identified as Chironomus hyperboreus, and one of which was identified as an unknown Chironomus sp. I then clicked “Tree” and selected “View Tree Result”, which returns a pdf with an auto-generated phylogeny of specimens in the database, based on barcode differences:

You can see the sequence I entered in red as “unknown specimen”. You can also scroll down and see the other specimens labeled as Chironomus hyperboreus which are included on the tree, for comparison:

Clicking on one of them lets you see who collected it, etc.: https://portal.boldsystems.org/record/ABCMB1613-23

This lets you at least see where the barcode falls among similar specimens, and what those are being called. Being able to determine how trustworthy those other IDs are, and whether the barcodes sort out nicely into different BINs for different species is something that takes some experience though. But that’s how you go about accessing the information if someone throws a sequence at you to defend an ID.

Paul,

I wonder if you might be interested in writing this up as a Tutorial?

I think teaching is already your day job, so it’s totally understandable if you use the iNat Forum to recharge, and you don’t want to feel like you’re still at work by teaching while you’re on the forum. But you know a lot, about a lot of things!