AI Suggestions in Fungal Observations

I apologize if this matter has been beleaguered elsewhere. At the risk of repetition, I’m posting it here.

I am a fungal curator on the site. To say I routinely come across thousands of misidentified observations, whose misidentifications can only be explained by iNaturalist’s rampant and wildly inaccurate AI-based suggestions, would be an understatement. I would conservatively estimate that upwards of 90% of all of iNat’s fungal observations – which, at the time of writing, amounts to 16.4 million individual records – are incorrect, and that most of these are the fault of the AI ID suggestion system. This garbage pile grows by the thousands on a daily, if not hourly basis, and to what end it’s not clear. The rate of error replication so dramatically outpaces the labor able to be performed by the few fungal specialists on the site. We’re not sweeping a dirt floor, we’re sweeping a constantly-replenishing dirt mountain.

To the outside observer, the primary incentive in the use of AI-assisted identification seems to be user satisfaction at the expense of database integrity. Contrary to what iNat’s AI system – and, consequently, many tens of thousands of individual fungal observations – might suggest, not every bunch of vaguely round, vaguely yellow, vaguely clustered, vaguely fungal growths on wood are Calycina citrina, but users will have this name thrust at them regardless, and go onto select it with some mixture of shoulder shrugging and pride, because even a bad identification is more psychologically satisfying than no identification.

We, the very, very few people willing to meaningfully, personally identify people’s fungal observations in our free time, would much, much rather see a seven- or eight-digit landfill of observations labeled simply as ‘Fungi,’ or ‘Agaricales,’ or ‘Ascomycota,’ in cases where the uploader is uncertain of a more precise identification. This would provide us with discrete piles to be worked through on an ongoing basis, as opposed to chasing behind a constantly ramifying rats’ nest of AI slop like a bull in a china shop.

This is not my first rodeo with large scale curatorial decision making on a biodiversity inventorying platform. I put a decade into Mushroom Observer before I ever came around to iNat, working hand in glove with that site’s developers over just as long a time period. There are important lessons about that site’s flaws/failures that iNat would be wise to learn from.

And yet, I have every reason to believe that this appeal will fall on deaf ears, if for no other reason than the fact that every other appeal I’ve made on these forums – all of which have been lower stakes and easier to implement than this one – have been abjectly refused/denied out of hand. If nothing else, let this post place “on the record” the situation described above, and present a space in which both the favorable and the detractors can weigh in.

I can attest to the fact iNaturalist’s CV for fungi is not simply insufficient, but wildly inaccurate. It once suggested to me that a photograph of mold on a cantaloupe was an earth star. There must be some way to mark observations as likely inaccurate en masse and reconfigure the CV to show more caution about fungal identifications. Perhaps it could be a new measure that requires at least two photos of things identified as fungi to be identified down to family, or, better yet, requiring photos from different angles or photos of the stalk as well as the cap from above, although this would probably be harder to implement.

Thanks for this post. I mostly dabble in identifying unknowns, and there are a lot of fungi. Since I don’t even have basic knowledge of fungi ID, I simply mark them as “Fungi.” I’m glad to have your confirmation that I’m taking the action that’s preferred by identifiers of fungi.

I very much agree with the main thrust of this topic – as an identifier of some corners of Ascomycota on iNaturalist, the situation with CV suggestions is fairly dire, particularly for some groups. However, I did want to highlight that the new iNaturalist (Next) app handles things in a much improved way, and frequently will suggest a higher-level taxon based apparently on strict consensus – most frequently in observations I’m looking at, Pezizomycotina, which is indeed usually the most reasonable ID. I think this shows that there’s the possibility for improvement.

Aside from the difficulty of implementing that kind of measure, fungi are so diverse that the requirements would be wildly different for different groups. None of the ones I identify have caps or stalks, for example; instead, for leaf parasites like powdery mildews, the identity of the host plant is critical. Some microfungi require microscopy as well at the very least. I don’t know of a good single solution as the problem of fungal observation computer vision suggestions is really a multitude of problems… and likely a suite of problems that are inevitable due to the development of the computer vision originally for macro-organisms like plants and animals.

A feature I’ve wanted for a while now is an auto-suggestion of the the lowest common clade that the top three CV suggestions go in, or for observations with multiple photos, an ‘average’ or ‘lowest common clade’ for every photo in the observation.

Aside from the difficulty of implementing that kind of measure, fungi are so diverse that the requirements would be wildly different for different groups.

Hmm. Good point. I mostly make coarse identifications of basidomycete fungi to the class agaricomycetes, so I was considering what might be helpful to that specifically. I have neither the programming nor fungal identification expertise to single-handedly design a perfect system. As you suggest, the solution would surely rely on multiple, complex measures.

Depends on who’s ears your trying to reach. You are not alone in dealing with CV issues. Many CV problems have been brought up before on other forum posts.

As it stands, there really isn’t any solution to this. For fungi, it is so widespread and out of control across many different taxon. I’m not even sure where you would start if you wanted to fix this. Perhaps recruiting a group of 50 volunteers? Its not just a couple problem taxa, its up to 100s. I have heard many snippets about certain fungi not even being identifiable by photo and needing microscopy, or other things like chemical tests. Fungi isnt my specialty though.

Unfortunately though the problem you describe is of such a large scale. It isn’t fixable by one person. The best solution i can come up with is that some taxon are complex and difficult enough, a Computer Vision algorithm shouldn’t be trying to assign species IDs.

Not sure exactly what your “appeal” is, other than a general complaint about the inadequacy of the iNat CV system for fungal identifications (which I am not doubting).

This might be one constructive suggestion to submit in the Feature Request category, perhaps to be limited to Fungi and other select groups where the CV usually falls short.

Another potential suggestion might be to simply disable CV for Fungi, and potentially for other groups with similar issues.

If anyone does decide to submit a Feature Request to address this issue, please (1) follow the template for submitting Feature Requests, and (2) provide a link to this Forum topic for context.

if - we would ID as ‘leaf parasite’ - there could be an obligatory … what is the host plant? And a prompt to make a linked obs FOR the host plant. Otherwise push the ID back to Fungi?

How about a default “Suggested by AI” mark, and then something like a checkbox for “I personnally identified this species”?

You can write - CV suggestion - as a comment. We already have that ‘twinkly stars / fairy dust’ icon for used CV. Does not mean - used CV have no idea what this is LOL - altho I have seen that as a comment.

I’d certainly love to always have the option of a higher-level ID, however high it needs to go, rather than (if it can’t come up with anything) just a bunch of different species. Though you might need the option of being able to choose the kingdom, since if you have both an insect and a plant present (say) it understandably can’t always decide which you care about. (For context, I only use the website. I think I’ve heard the app may be better for this?)

I finally decided to create a profile on this forum just so that I could respond to this because it bothers me SO MUCH! My first question is, how can AI possibly get better at IDing mushrooms if the vast majority of online photos are incorrectly IDed??? The original poster very eloquently described the problem and how big it is. I think two simple changes would GREATLY help this issue:

  1. AI should NOT be allowed to ID mushrooms to species level. Most mycologists have a hard time IDing to species level without data such as microscopy work, smell, etc. So why would we let AI guess at a species level on an observation with one crappy photo of the cap???
  2. As marilyngx suggested, there should be a REQUIRED mark that states “suggested by AI” on any mushrooms that weren’t IDed by the human. This would help solve my problem of: “Does this ID seem very incorrect because the IDer knows nothing about mycology and just clicked the first suggestion, or, is this a fungal expert that knows something I do not and correctly IDed a less-than-optimal observation?”

On In the browser, there’s already a little icon (the chevron sheild with the sparkles) indicating someone selected an ID off the computer vision.

If the CV had a thing on the page to upload observations where you can type in the domain/kingdom it’s supposed to be to make the CV suggest IDs in that domain, that would be helpful.

I think the point is that there’s no way of telling whether someone knew the species and just picked it from the dropdown to save typing - and typos! - (as I do if I can), or had no idea and just picked whatever came out top.

I agree, as long as it’s optional - because if it was required, that would add an annoying extra step in the majority of cases where there’s no real question.

I feel your pain. I don’t think the issue is falling on deaf ears, rather that it is not easy to fix. Hopefully the changes to the new app will help the situation somewhat, but it won’t be a complete solution.

It seems useful to link back to the observation accuracy experiments done some time ago. These indeed find that fungi is the iconic taxon with the highest % of misidentifications on the site, at least when looking at Research Grade observations (Chromista and Protozoa are worse when including all observations). That said, in the sample selected for the experiment, 86% of Research Grade fungus observations were assessed as accurate. When looking at all observations, the sample shows only 72% of fungal observations assessed as correct. Better than 10%, but still far from ideal.

One thought this gives me is that it may be worthwhile, and more satisfying, for fungus identifiers to focus more on fixing incorrect Research Grade observations, rather than doing battle with the flood of incorrect Needs ID observations.

Other than that, I’m not seeing a clear, actionable solution to this issue that doesn’t come with major costs or downsides. One could have some system for manually designating taxa where the Computer Vision should never suggest anything finer than Genus, or even Family. For example, whenever the Computer Vision “thinks” an observation is Russula emetica it could just suggest Russula sp… but who decides that, and how? Would there be a voting system? Open to anyone, or just curators? How would that information be integrated into the way the Computer Vision works? What happens with taxon changes or when users disagree? And so on. It wouldn’t be a trivial change to implement, but I agree that something like this is needed, and the details need to be hashed out to get to the point of making it a feature request.

As much as I (occasionally) complain about fungal identification in regards to the CV, I have to admit that it isn’t always bad. From my experience, actually, there are some groups that it is quite good at - namely, iconic, easily recognizeable, common taxa, also known as ‘exactly the groups that we don’t want experts wasting their time on anyway.’

I’m my area, I mean things like Trametes versicolor, Laetiporus sulphureus, Desarmillarea caespitosa, Megacollybia rodmanii, Concocybe apala, Leucoagaricus americanus, Lacrymaria lacrymabunda etc… tend to be very accurate. They have obvious features and there are LOT of observations of them.

Even less obvious-at-first glance species it does get right - https://www.inaturalist.org/observations/290749153 this observation, for example, the CV IDed as Homophron spadiceum. You can see that I forced it to genus level because I wasn’t fully confident in that id, but micro confirmed that ID, and I have no doubt that if I do sequence it (I may not bother), the sequence ID will come back as a confirming match.

Here is another example that I already uploaded, but replicated here for the sake of a visual aid; it is indeed very accurate for this fungi (It is, indeed, Oudemansiella furfuracea, without any doubt,) but it at least shows some uncertainty because the first photo is lacking in a few details that the subsequent photos clarify.

The problems really start to arise when we either get into cryptic groups, groups that lack properly study and clarity (so that identifiers are unsure of the true macromorphologic variability present in the mushroom and they taint the CV -Laccaria laccata is one example of this), or very old genera that have implanted themselves in the mind of the general public as The Name for a group and thus end up being kind of a dump species for anything that looks vaguely like the original description (Russula emetica is a good example of this)

Here is an amanita I found recently - a good example of both a cryptic group and one that tends to get used as a dump genus. A. bisporigera is the top option - and this is certainly an option, but there are several other white amanita in Sect. Phalloideae in my area that this could be. A. suballiacea, A. amerivirosa, and maybe even A. magnivelaris and A. ellipotosperma are options - plus there is a well-understood nom. prov. by the name of A. “sturgeonii” which is actually one of the more id-able-by-just-macro members of this group but is not yet described and thus will never, ever show up as a suggestion. A. bisporigera tends to be the one that folks default to, so everything gets dumped there.

Phalloideae isn’t even the most egregious example here, this particular mushroom would be easy to figure out with just some basic micro (amanita are, for the most part, pretty well understood in the United States, compared to some other groups.) If I had used a red russula as an example, the results wouldn’t even be as close as you see here.

Another issue you run into is folks not keeping track of new publications (which, understandable) who stubbornly stick to older names and gum up the system, confusing the cv

This is Macrolepiota macilenta; this species was described in April of last year with another Macrolepiota species, M. pallida - finally giving name to the eastern north american species of Macrolepiota that have been known about in sequencing for a while but not formalized via description.

M. clelandii is a New Zealand/Australian species, and M. procera is an old European genus that has ended up as the dump species for the genus for a long time.

Now, I deliberatly left the location out of this example because I suspect some people are doing this and not adding locations until after they have used CV. With location, the suggestion looks more like this

Much more accurate, right? As I said, I suspect some folks just don’t add locations until after they’ve gotten the CV idenfication, because I am still running into CV-suggested observations for Macrolepiota procera in the Eastern US, despite me periodically going through Macrolepiota observations in the US and cleaning them up entirely. If you go look right now, there are two CV-suggested observations of M. procera in the eastern US RIGHT NOW.

No amount of extra safeguards is going to help if people aren’t utilizing the tools properly.

I actually don’t think the CV is as much of a problem for fungi - at list it’s not any more of a problem than any other group that has cryptic or poorly understood species that require specialist knowledge and tools to identify. I think the problem, ultimately, comes down to our broader cultural relationship with fungi, especially here in North America - there simply aren’t that many people, in the grand scheme, that give a hoot about truly accurate fungi identifications. Even many people that are IN to fungi as a hobby engage with them only as far as they are useful to them, IE, can they be eaten, used medicinally, or used as a drug.

This isn’t Aves or Lepidoptera; there aren’t MILLIONS of people who observe these as a hobby. The ratio of people able to make accurate fungi ids to the people submitting them is just too skewed to really get some of the worse groups cleaned up.

EDIT: FWIW I’m basically at the point where if I need to get a better idea of the true range maps of a species, I only look at observations that include the ‘DNA Barcode ITS’ field - ex https://www.inaturalist.org/observations?subview=map&taxon_id=125390&verifiable=any&field:DNA%20Barcode%20ITS= . It is niche, of course, and it is a minority of observations that have sequencing, but it shows observations that have molecular data to back up the ID. Most of us are pretty good at kicking observations back to genus when there is zero chance of it being that species sensu stricto.

I’ll agree with the “very, very few” part. As someone who has tried to use the identification resources available to me as a nonspecialist, only to have my fungal observations knocked back with comments to the effect that we do not and cannot know what it is – well, I have a term for that, but my posts have been flagged in the past for using it.

If you had simply ended the sentence right there, that would reflect what it looks like to me as an observer. I used to get out my mushroom guide or look for online information to provide the best initial ID that I could, only to have my efforts brushed aside. Honestly, it doesn’t feel worth it to put in more effort than it takes to trust the CV, since the outcome is the same regardless.

I have read the journal entry on how to photograph mushrooms – the different angles to take, etc. I have added notes about the taste and about what trees (if any) are nearby as possible mycorrhizal partners. And still only 43 of the 111 observations that are at complex level or below – just 39% – are at RG. And that’s just the ones at complex level or below; I have 198 fungi observations altogether, which means that 43% (87/198) are above the complex level – more than the number of my RGs. This does not feel like a very good rate of return on putting in an effort.

When observing boletes, I use The Bolete Filter. And yet, of my 11 Boletaceae, 6 are at the species level and ONE is at RG. Again I say, this does not feel like a very good rate of return on putting in an effort. Would the stats be any different if I just trusted the CV?

If by “the identification resources available to me as a nonspecialist” you’re still referring to just Googling things, the fairly negative reaction to that probably has a lot to do with the fact that there’s a lot of bad information out there. It probably also depends how old your mushroom guide is, in much the same way.

Meanwhile, the complaint about how few research grade observations you personally have seems to be indicative more of a lack of additional experienced fungi identifiers to deal with the giant mountainous pile of observations that need identification. Yours are part of an immense number of records for which there aren’t ever enough experts to keep up. Just because your well-photographed organism isn’t being identified right away definitely does not mean you should completely stop “putting in an effort”. You are relying on the kindness and attention of people volunteering their time to comb through your photos; it seems to me that giving a little more of your own time to make it actually possible to identify your observations is the least that can be done. I’m not even sure what you’re suggesting as an alternative – no longer taking multiple angles’ worth of photos? Blindly accepting CV suggestions even when the Computer Vision output itself says it’s “not confident”? The question is why this is so important – do you actually outright need your observations to be Research Grade for some purpose? I ask because I’m genuinely curious what makes your percentages so critical – but also because I’m one of those identifiers working extra hard to bring updated taxonomic treatment to regular observers, and have never seen this level of frustration from others. I’m still working on the field guide I mentioned in that other topic linked above, as it so happens – but that too will not be available at everyone’s fingertips right away. Effort is a necessary part of identifying and observing more difficult groups of organisms; I can imagine no reason why getting identifications from others for free would be effortless.

Nothing that I can think of will solve the overall problem of misidentified fungi, but there are some things you can do to get your own good observations to Research Grade. Check out who seems to be doing good identifications for your area. Identify some of their observations. Then ask if they would identify some of yours. Hopefully some will “follow” you, which should improve your the ID rate for your observations.