What is this - iNaturalist and generative AI?

Thanks for highlighting the changes in phrasing that have been used. It seems worrying that “through technology” has been added. Let’s hope it’s an unfortunate choice of words indeed and not a shift in priorities. From the same About page, iNat is described as:

iNaturalist is an online social network of people sharing biodiversity information to help each other learn about nature

If iNat’s goal is for people to help each other learn about nature, then I think teaching others about identifying species is best done through something like this: Create an “identification center” with guides, curated subcategories, events, and more. This enhances the community itself, it makes people learn directly from each other, and it seems well-aligned with iNat’s mission.

If iNat’s goal has indeed been changed to achieve this “through technology”, then I think using the latest, shiniest technology is in alignment with the objectives. However, I think that clarifying whether or not there is an intentional strategic change that gives technology a more prominent role is warranted.

I know there is temptation to be suspicious of any inconsistency in wording, but in the end this wording only describes what iNaturalist has been doing all along, in full view of the community, and I choose not to look for anything nefarious in it.

Me too.
The backlash to this is crazy.
Most comments against it seem wildly illogical or informed by tabloid headlines / Instagram memes rather than hard data.

There is no proof of this until it is trialled. It really depends on the implementation.
There is no reason to think it couldn’t help with the opposite.

Re: hallucination.
People are commenting around this as if the existing CV doesn’t hallucinate about species suggestions at times too. I’m genuinely confused. Would those who are against use of genAI here also like to remove the use of CV entirely? The CV causes loads of problems, but on balance it’s a far more effective tool than not. For me it’s one of the reasons I use iNaturalist. What reason is there to think a genAI implementation wouldn’t be an overall net positive?
Especially given there has been no trial or demo to even give a basic idea of functionality?

This is how I felt yesterday. The anti-LLM arguments are so pervasive atm in social media.. it’s exhausting to try and counter such emotionally-heated / baseless responses. I would rather be out in nature recording!

iNaturalist has been using AI since forever.
This makes no sense.
Do you never use the autosuggest? Or you are against the existing use of CV?
If not, why would you think this is any different?
At least environmentally or in terms of hallucination, there is no reason to think it will be.

Still catching up on - while I was sleeping - but I MUCH prefer to have a single string of comments. And to have them here in the forum - where it is better formatted. And threaded!
Thank you iNat.

Well! This is a hot topic! One only has to add the letters AI to make people passionate about any topic… :grinning_face:
I’m sure similar reactions occurred when transitioning from paper to internet sources and so on.

I think it’s a valid use and worth exploring, and financed by a grant, so also a responsible use of iNat resources. I’m curious to see what the result will turn up to be. An imperfect gadget tool.

I also think the staff is aware that the real value of iNat is the people willing to give their time and expertise to identify species and… on top of that, provide useful tips on identifying them. This is not only better than any gen. AI will be, but also so much better than any of my fields guide. As long as this remain at the core…

I can also threat to leave iNat if they don’t like all my posts, but I believe this isn’t intellectually the best approach to make a valid point about the topic.

The CV is an image pattern-matching technology. It does not make any claim beyond “based on our pattern matching technology, we think this image most closely matches this other image”. It does not hallucinate false statements, because it does not generate text.

As others have pointed out, if the proposed “AI” in question was simply a comment aggregator that pointed to where other people have made helpful ID tips, basically no one would have a problem with it. It’s the part where the AI is generative, meaning it attempts to synthesize other text to form its own sentences, where the danger lies and the hallucinations begin.

I would say it’s because of the negative impacts genAI has had in every other aspect of their life. That’s probably the reason.

I oppose and condemn iNat’s project to “improve” species suggestions using generative AI. This project - funded by, advised by, and open to being infrastructurally supported by Google - is a mistake.

For those familiar with species identification and AI tools, it should be obvious that trying to use an LLM to explain the reasoning behind a species ID is a blatantly silly idea. Without constant verification by expert IDers, it will be untrustworthy, misleading, and unaccountable. For it to be usable, it would take so much AI-babysitting work from expert IDers that it would have been much easier, simpler, more reliable, and more community-focused for those expert IDers to have built a user-edited wiki in the first place.

While the iNat computer vision is an amazing tool for science, generative AI (a very different technology) is, as iNat plans to implement it, decidedly not. In fact, it is anti-science: it stands to make community science worse.

If I want to learn why my observation is what it is, I want to hear from a person. A person has reasons, and can articulate those reasons. A person can explain why. I want to hear a taxonomist’s nuance, an amateur’s passion, an expert’s tip. Most importantly, a person can cite their sources. A person can admit when they aren’t sure, when it’s a “most likely this” kind of situation, when it’s more vibes and gestalt than professionally keyed out.

genAI can do none of this, simply because genAI does not have reasons and cannot articulate reasons. Forcibly training a glorified autocomplete, until it appears like it can sort of articulate reasons, most of the time, is an extremely inefficient and alienating use of your community of IDers.

iNat is precious to me. I love it. It has woven webs of trust, balanced accessibility with expertise, and built a fun, empowering ecosystem of learning.

It is this sense of community that is iNat’s greatest asset. It is the very thing that makes it unique. And it is exactly this that is threatened by its genAI project.

We didn’t ask for this. We don’t want it. Please get rid of it.

Many have voiced concerns with genAI that I consider broader and more structural in nature. Some of these include data privacy concerns, intellectual property / labor concerns, environmental concerns, and concerns related to Google, including its well-known business model of allowing small orgs to use their compute power while eventually making them dependent on Google computing services.

I believe many of these wider, systemic concerns are, to varying degrees, valid reasons to condemn this use of GenAI.

My main point is that even if none of these concerns existed, I still wouldn’t support this use of GenAI. Even if training the LLM had zero carbon costs and was funded by another company, I would still condemn it, simply because it can do nothing but make community science worse.

Even if you haven’t seen what the AI tech bubble has been doing to our world, you should still oppose this project, for the utterly pragmatic reason that a fancy autocomplete does not and cannot reason about why a particular observation was identified to its particular taxon, nor can it articulate such a reason.

P.S. I’d happily contribute my time and expertise to a user-edited wiki or similar community-driven project to improve iNat species suggestions.

P.P.S. the one use case I can envision for LLMs on iNat is as a functionality of an iNat-internal search engine that helps users easily search iNat obs and pages to find comments, notes, ID explanations, etc. that are relevant to their search keywords. It is in this kind of search engine role - as a circumscribed library-research-type tool - that a glorified autocomplete may not be completely useless. Of course, the AI slop vision laid out by the iNat team in their blog post has little to do with this idea. This hypothetical use case for genAI on iNat is also not immune to the broader structural concerns mentioned earlier.

For all the iNatters who say they are ‘eager to help with a Wiki’
https://forum.inaturalist.org/t/ways-to-help-improve-inaturalist-taxon-pages-through-wikipedia/2680
https://forum.inaturalist.org/t/creating-missing-wikipedia-articles-for-inats-observations-of-the-week/18057
https://forum.inaturalist.org/t/needed-images-list-wiki-for-existing-wikipedia-articles/18874
https://forum.inaturalist.org/t/inat-to-wikipedia-pipeline-thank-you-and-tool-tutorial/49550
https://forum.inaturalist.org/t/help-the-world-out-and-create-this-page-on-wikipedia/6101

What are you waiting for? The need is great. Go for it.

After some time to reflect, I wanted to chime in again share some thoughts that may further clarify both my own stance, and possibly cast some extra light on why I think this backlash has been so strong.

I understand that a lot of the public response to this has probably seemed a bit kneejerk, vitriolic, and even confused. I don’t envy you, and I do genuinely apologise for the abuse I’ve seen directed at the iNat staff on other platforms. If the idea of seeking to incorporate new technological advancements into your app and keeping pace with global tech has been part of the ethos of iNat from day one, then I can understand why a backlash of this magnitude might have truly come as a surprise.

The project you’ve described has been a fairly simple and perhaps even logical expansion of the computer vision that’s already part of the app. And I believe you have described it transparently and been open with your intentions for the project. But this hasn’t done much to calm fears and anger from much of the userbase because, at the end of the day, this issue represents a critical disconnect between the philosophical axioms of the staff and much the userbase.

Generative AI has caused an enormous social upheaval since it’s proliferation, and created unanswered and often ignored questions that threaten the livelihoods of people in all fields. Outside of science, I personally am an artist and 2D animator, and in the past few years I have had to grapple with the fact that much of society has embraced an automated system that plagiarises work like mine and spits out an empty aesthetic recreation of it, largely for the purpose of not having to deal with the nuisance of people like me. Technology that allows a computer to generate and spit out all our culture for us so we can all get to work somewhere useful, like an Amazon fulfilment warehouse. The world calls this “progress”.

Finding factual information has been made harder by pages and pages of hallucinated AI slop and deliberate misinformation. Connecting with other people is made harder by not even knowing if everyone online is real. There are a slew of very valid environmental concerns, privacy concerns, danger to public access to information. Things have become worse for a lot of people because of this technology, and it’s becoming hard to escape it.

So when a large portion of the iNat userbase riots over the development of new GenAI with Google, they aren’t doing it because they’re worried about the specifics of how the AI is going to be implemented. They’re rioting because genAI is being added at all. The arguments made by many about things like environmental impact, image scraping, etc. might not be super relevant to this specific project, and thus they probably aren’t great arguments, but they nonetheless represent broader, real frustrations people have with genAI as a whole. Even if in isolation a specific use of genAI is largely harmless or even somewhat beneficial, many people don’t want to support apps that use genAI on principle, or see their money go towards developing it. And I agree with them.

Perhaps this project will go ahead despite the backlash (the grant has already been accepted after all). And maybe all it will end up being is a small, unobtrusive generated ID suggestion of dubious value and accuracy that can be safely ignored or even turned off. And that would be nice. But even if that’s all that happens, I think the implementation of genAI like this is, for me and many others, ultimately a kind of crossing of the Rubicon. It would be a sign to many here that iNaturalist is not the little oasis of solid, trustworthy human interaction and expertise that so, so much of the us have come to love and support iNat for being. It would push away lots of people who would give, or even have given, so much to this app, because it would be a sign that the values of iNat do not align with an enormous portion of its userbase.

I don’t want to leave iNat. I have sung the praises of this app for four years now, and I still care deeply about the goal of citizen science. But if this is the unavoidable direction things are going, then regrettably I feel that I could no longer support iNat through donations and advocacy. And as much as I want to stay, and hold onto the parts of this app that I still love, I may personally just have to say “thank you for everything you did and everything you were up until this point, and good luck”, and then move on from iNat for good. I know I’m not the only person who feels this way.

I can’t comment for everyone but “AI” never really has been a single thing and I believe it is reasonable to accept the computer vision applications but be very sceptical of this proposal.

Right now, “AI” is often an umbrella marketing term to encourage us to take any past positive experience with machine-learning solutions like computer vision, text generation, automated translation, etc. and what we know of really specialised machine applications like protein folding and assume that it carries over to large language models and all the chatbots, agents and other tools that are built on top of them.

Eliding these differences is dangerous. Species identificaiton using computer vision is amazing, but it works because it is not hard for iNaturalist to construct well-annotated training datasets for this purpose, and the domain of the expected responses is well-defined and closed (one class for each taxon represented by input images). Using all iNaturalist’s observations to predict the species likely in a given area is another statistically reasonably grounded application of ML.

With suitable constraints, LLMs can be made reasonably reliable for many use cases, but 1) the training set for an iNat identification text tool will generally be small for each species, highly specific to articulating what can be seen in a particular observation’s images, including a mix of taxon-specific technical language and different simplified descriptions, etc., 2) the foundational LLMs will not have much starting leverage on the specifics of these species-by-species cases, and 3) the range of possible model outputs is infinite, with desired, valid and helpful responses only representing a tiny sparse subset of that range and with no obvious common patterns that can be plugged with specific terms harvested from the available comments.

I am sure this could be made to work in some cases. I am equally sure it will deliver misleadingly confident and partially correct but ultimately unhelpful answers in many others. I am also sure that, even if scaling is a challenge, a well-designed tool for the community to edit text elements and to annotate images (like the frog example) would allow knowledgable curators to knock this challenge out of the park.

I really hope iNaturalist can be clear and precise on how “AI” and ML are expected to fit in, and that any seriously pursued applications are adequately grounded in information theory and statistical plausibility.

Rereading the blog post
By providing explanations in addition to a list of suggestions, iNaturalist hopes to more effectively grow a skilled community of naturalists who have the information and tools to improve and vet the data on iNaturalist.
Identifiers are drowning in a sea of obs. We need to mentor more identifiers. This looks promising.

Can we distinguish between predictive AI (like iNat CV) and generative AI, which have many more dark side than bright side?
Also deleting their own account is their rights and choice, not yours or iNat.

We are waiting for people asserting that all it takes is putting ID tips and advice on some third party, world-editable, unrelated website with a different set of rules and aims, to 1/ read (again) the many comments (here, in the blog, on BlueSky) explaining why this is not the best idea and 2/ realize that people who contribute to Wiki with scientific data and expertise for longer than iNat exists may have already commented on that.

So, read again, and reflect on that. Go. What are you waiting for? ;)

I’ve thought of something like that too. Why not use a user-curated wiki as training data for the AI? That way, people can always also check a human curated source, and some of the worries about people’s data being used to train an AI model could be eased, as nothing apart from the Wiki would actually be used as training data. Don’t want to contribute to the AI? Don’t contribute to the wiki. Everything else is “safe”.

Also, the edit to the blog-post has reassured me somewhat. I think using GenAI to give tips such as “photograph from this angle too” or “look at this feature for identification” is a MUCH BETTER implementation than it giving an actual explanation and reasoning for a CV-ID. I think the observer should still have to decide what they see.

An AI comment like “Look at the scutellum, is it black (sp. 1) or white (sp.2)?” means, the observer needs to check the images themselves, and can only make an ID if they know what the relevant feature is.

An AI comment like “Sp. 2 because scutellum is white” would lead to the observer just blindly agreeing… No idea what a scutellum is or whether you can see it on my photos, but if the AI says so, it probably is. ⇒ Overconfident ID

Not sure what others are suggesting, but I’ve taken references here to a “wiki” not as references to a Wikimedia site but rather to an interface inside iNaturalist that would allow curators to enter simple text in a structured way so it can be injected into the automated id processes, etc. In other words, an interface that (for the frog example) allowed this information to be entered (using the same species selection interface as in Explore, etc.):

Ideally there could also be a geographic scope for applicability of the characters as a useful tool and annotated images like the frog ones.

All of this could be entered through a simple UI in a way that kept it all completely structured.

When the CV offers someone the suggestion that their frog is one of these species (at least within the overlap range or the specified geographic scope), they would also be offered this information and asked to check that they can see the specified characters.

Curators who today repeatedly fix such pair-wise misidentifications could then inject their expertise at the point where most mistakes occur.

Identification sets could include many more than two species where that makes sense.

iNaturalist could provide such a form for curators and ideally a way to select suitable species images and attach labeled characters to the appropriate point in each image (as with the triangles in the frog images).

That seems eminently workable to me and 100% in line with current iNaturalist culture.

Hi Donald, I very much agree with what you’re saying. Plus there are novel approaches how characters tables can be filled by a community, which are then used to generate identification pathways on the fly. You may want to look into the references I provided here. Cheers!

What I was referring to: various users have been suggesting/requesting an on-site wiki of sorts, where iNatters could contribute ID tips for iNat taxa, using iNat contents; it has been repelled by suggesting people go and edit Wikipedia.org pages instead (ignoring previous arguments against this “solution”). It also came with a few kind/well-meaning comments, in the line of “if you haven’t contributed your ID tips on Wikipedia.org by now, what makes you think you would share them on an iNaturalist wiki in the future ?”

Please flag derisive comments rather than engaging with them. Also remember to assume people mean well, even if their comments could be read in poor taste.

This is indeed the standard way the term ‘wiki’ is used in a context such as this. Not really sure why people would be thinking it refers to wikipedia specifically.

Sorry, I don’t really see a difference in how this will impact users making wrong IDs.
Hallucination of false IDs is no different to hallucination of false statements.

From a user-perspective, you will have those who just click on an option in the auto-suggest regardless, and those who use it more carefully and choose a higher level ID, those who tag in expertise to double check, those who dig directly into sources themselves… etc, etc.

If I upload a hoverfly in India and it has additional information to support the auto-suggest I don’t see that this will cause innately more problems than the existing autosuggest does.
Especially if the implementation involves the option for users to flag up if something is problematic.

With ChatGPT you can ask for a source for any information that it is pulling from and it will either point you toward a made-up or weak source…or point you correctly to a genuine source. Again, if implemented well, this could be no different.

I would love to be able to instantly have a suggestion for a diagnostic character on a species I don’t know. I would love to be able to have a link to a source for the diagnostic character. Trying to dig this information out of the existing data set is impossible as a user…having to tag people in every time to find out feels suboptimal… and being tagged in on an observation where you have to tell people for the 100th time that you can’t go to species on this one without microscopy is beyond tiresome.

That’s a bit broad to be meaningful here.

It seems to me this will really depend on implementation which in turn will depend on trial and error. But again, I don’t see this as likely to be worse than the existing autosuggest in problems it will cause in the long run. I think it is unlikely it will offer a net negative impact.

So I guess I just agree to disagree with all the naysayers.

But either way you must be having a different experience of genAI to me. I am yet to see a particularly negative impact of genAI on my life. The opposite actually. I find LLMs to be fantastically useful tools. I find iNat’s CV to be a fantastically useful tool.