What is this - iNaturalist and generative AI?

I’m actually more disturbed by the knee-jerk reactions by so many on this forum than by the possibly poor rollout of information about this project by iNat staff. I expected a little more patience, consideration, and rational thought by iNatters who are science-oriented people. But I guess I was mistaken.

This is very good news, and alleviates my most major fear. I wish some…any…of this had been communicated, and I am glad you have contacts that you can get these answers for us. It is sad though that it has been done through an intermediary! Maybe…iNat should hire a communicator (only half joking…)

[quote=“graysquirrel, post:431, topic:66140”] A small amount will be going to some specific grant-related projects, which, again, are not actually genAI.
[/quote]
So…why are they calling it GenAI? I’m not super techy, but I understand the basic differences.(and for example, someone else asked “why should it not be called AI” well that’s because it isn’t actually artifical intelligence, that’s just been the term decided to apply that everyone has gone with but I digress and cannot find the link explaining why LLMs are not actually AI, etc etc)

“AI” note additionally - I have finally downloaded the new phone ap today, and notice it is called “AI” instead of “CV” now. I find that super disheartening and may go back to the old ap just because of that tbh.

I’d be curious how many folk actually left. I was among the I will delete my account crowd at the start, due to the potential for my data to be handed over to Google. Obviously, I have not done that yet, as I am intelligent enough to wait for some answers to the privacy questions I outlined very early on (as were others). I realise I have a large build up on here (over 3600 obs, and almost 24K ID’s) and would not dump that without confirmation of likelihood of that scenario.

Thank you! Having an opt-out is probably for me the most singular important thing so I can choose my use of it, and choose if I want my comments included, etc.

As to everyone saying they feel they can trust the staff even less now, welcome to how some us felt after June 2023. That trust has still never been fully repaired for me. I do use iNat less now, and more “for my own gain” (i.e. when it is helpful to me, and helping those I follow and have formed relationships with, rather than giving back in my ID time widely) and until this storm settles it will probably remain that way. I mention this because I want people to understand “shits happened before”, and how those of us who felt extremely wronged, and confused, have handled it. A few did leave, but I think only two did in the end. Most us involved dialed back, at least for a while. Some continued on as usual. There isn’t a need to think of this situation as a binary thing of remain as usual, or delete it all. There are so may 3rd options!

I’d encourage you to reconsider your perspective here. There are two things here, the staff and the community as a whole. While you may be directing your “they” at the staff, “they” can also be defined as “us” or “we”. The question should be framed as, can we “afford losing a few hotheaded” identifiers and observers? The observations added by anyone who has deleted their account are no longer available for people like me to learn from. Was it a significant loss? I’d have to see their stats to know for certain, but I want to stress what it means to the community when someone deletes their account. The person not only hurts the staff, but the rest of the community too.

Thank you for continuing to engage with us. Most of these reassurances are great to hear, especially the reiteration that it will be opt-out.

Other points made in this post continue to be serious causes for concern. The first one is this:

Google.org isn’t expecting anything from this grant except for us to explore Gen AI technology in the context of our work to come up with solutions for how to better surface and organize expertise on the site.”

This quote implies that the expectation of the grant is, still, to fund GenAI tools to help collect, organize, and share the expertise of iNat users.

I understand that what these GenAI tools will look like is very much unclear. This makes sense. But, again, the latest official proposal from iNat is, still, a mockup of an AI-generated text description that tries to explain why an observation has the species ID that it does.

So long as this implementation of GenAI remains a likely or even possible course of action, I must continue to oppose it.

Another point from this post compounds these worries:

“Gen AI means so many things to different people, but to us it probably means that the demo we will make will leverage interacting with an LLM in some way.”

To me, this quote unfortunately reinforces the official announcement. It tells me that iNat is still leaning towards using LLMs to generate natural-language sentences to explain species IDs on observations, as specified in the blog post.

Once again, I urge iNat to make another public announcement to clarify the situation and address, explicitly, whether the GenAI text description approach specified in the blog post is still the most likely course of action.

Okay so within the last few hours we’ve been told that it isn’t going to be genAI:

Then we’ve been told that it actually is going to be genAI:

And yesterday we were told it would potentially be something more like collecting relevant comments and putting them together without generating anything new, which isn’t genAI:

The blog post called it genAI and talked about trying to implement a feature that a previous blog post said involved Vision Language Models, but this latest clarification seems to imply that’s not actually what’s happening (unless I’m misunderstanding, in which case, my apologies for that), so if that’s not happening then why did the blog post focus on that as the main feature that’s being developed?

I appreciate the apology for the poor communication at least, and the fact the staff have been responding here to try and clear things up, but now I’m left wondering which of the clarifications we’ve had are actually the correct information and why there are so many inconsistencies in what we’re being told.

is a hamburger a sandwich?
is a tomato a fruit or a vegetable?
is my friend who eats fish a vegetarian?

if you’re concerned about something when it’s labeled as Gen AI but you’re not concerned about the same thing when it’s not labeled as Gen AI, then i think what you should call it depends on what sort of state of mind you want to have for yourself.

if you have concerns regardless of how it’s labeled, then it’s still worthwhile to express what those concerns are, if you haven’t already (ex. energy use, loss of privacy, misinformation, devaluing of humans, etc.), and iNaturalist may or may not be able do things to address those concerns (ex. energy audit, ability to opt out, feedback system for results, transparent sourcing of information, etc.)

I’m not entirely sure how you came to the conclusion that my concerns are based on what the AI is labelled as, especially the idea that I wouldn’t be concerned about generative AI if it was just given a different name.

In this comment I was trying to express my concerns about the conflicting information and lack of transparency about the situation, but I am also concerned about the use of generative AI because of the unethical training methods (using data from around the internet without the consent of the people who created that data), environmental issues, and potential for spreading misinformation.

I didn’t feel it necessary to bring up the same points other people have made again and again in this thread when I made that comment, but I can assure you I’d be concerned regardless of whether they said genAI or LLMs or just gave a description of wanting to scrape comments to generate information.

The description Tiwane gave of something that would collect relevant comments that could then be sorted through and put onto a page so that useful information is all together sounds like something that wouldn’t need to be trained the same way something like an LLM would, wouldn’t have the same risks of generating incorrect information, would presumably have all the gathered comments be credited to the people who wrote them, and most likely would have similar energy costs to the CV, rather than the kind associated with generative AI models.

If an LLM is being used, however, that would come with all those concerns I mentioned before. You could call generative AI something else and I’d still be concerned about those things.

I meant the staff, from a numerical standpoint.
Regarding the disgruntled observers who have left during this epic PR fail, from what I gather on social networks, some may come back, and most if not all seem keen on engaging with nature and sharing stuff… somewhere else. Beyond iNaturalist there are many other communities and platforms, suitable for naturalists of all venues. This is good news, diversity is nice to have. No one died, potential experts and identifiers are still alive and kicking somewhere. A few obs may be lost forever on here, but hey! many more will be (re-)uploaded here or there, soon.

(Incidentally, it is nice how iNat is still free from the “vendor lock-in” issue: even if not straightforward, thanks to .csv exports and the API and the sweet guidance by @pisum it is feasible to move away without forfeiting data/metadata)

Apologies if this is off topic, I can take it to DMs if so. But what is this referring to? An online search yields something about AI image scraping, but I’m not sure how that would erode trust in iNat since scrapers are notoriously difficult to fend off.

It is very off topic, if users have questions they can DM me. I simply wished to include the concept that “iNat staff have messed up before” and “we figured out how to carry on without some great science harm, or harm to our own ethics” to encourage people to think beyond the binary all-in and total deletion.

i don’t see much inconsistency between “scrap[ing] comments to generate information” and “collect[ing] relevant comments” (emphasis mine).

suppose we have this these identification notes related to Rudbeckia hirta:

  • “has hairy green parts”
  • “has hairy stems and leaves”
  • “it’s fuzzy”
  • “leaf faces and stems hispid to hirsute”
  • “tallos y hojas peludas”
  • “hirta = hairy”
  • (1000 more variations of the above)

do you want a page that displays 1006 items that say the same thing in different ways?

if not, then as far as i know, any automated method of trying to present that information in a meaningful, more concise way is going to involve some sort of language model.

for me, something like this would be ideal:
“community identifiers note that Rudbeckia hirta has hairy stems and leaves (hirta = hairy). more precisely, the leaf faces are stems are hispid to hirsute. observations 1, 2, 3, and 4 have discussions involving identifiers X, Y, Z that may provide more insight.”

but even if the system did something like grouping everything together into logical conceptual groupings, i think that still involves a language model to know how to form the logical groupings (maybe presented within, say, an expandable list). for example:

hairy stems and leaves >>

  • “leaf faces and stems hispid to hirsute” (ex. obs 1 with identifier X)
  • “hirta = hairy” (ex. obs 3 with identifier Z)
  • etc…

I read zero posts in this thread, but here’s what the AI summary just generated from 416 posts:

The iNaturalist community reacted strongly negatively to a newly announced grant from Google to explore using generative AI to improve species identification suggestions, with many users threatening to leave the platform.

Concerns centered on potential inaccuracies, the environmental impact of AI, a distrust of Google, and a fear that AI would undermine the valuable human expertise and community interaction that define iNaturalist.

graysquirrel initially led the opposition, compiling data showing overwhelming community disapproval (#373). However, after a three-hour meeting with iNaturalist staff (@loarie and others), graysquirrel and procyonloiter reported a positive outcome, clarifying the situation and alleviating concerns.

The staff emphasized that the grant is a standard funding source, that no user data will be shared with Google, and that the AI exploration is a small part of the grant’s purpose, focused on surfacing and organizing existing expertise, not replacing it. The project will be a demo and will not be implemented if it proves unhelpful.

loarie apologized for poor communication and reassured the community that iNaturalist’s core values and community-driven approach will not change. Many users, including Masebrock and holocene_matt, expressed relief and appreciation for the clarification, while others, like spidercat and dlevitis, remained cautiously optimistic.

Frankly, iNaturalist should probably reconsider using generative AI anywhere on the platform, including this forum.

This website is putting words into my mouth that I did not say. Are we okay with that?

Not sure if iNat has any control over this feature of the Discourse platform. But so it doesn’t get lost in this ocean of a topic, I suggest posting your request in a new topic under Forum Feedback, to prompt staff to investigate the options at least.

Yep. Hotheaded reactions with permanent, damaging consequences based on utterly incomplete information. Not an edifying thing to see.

It’s one thing to pause uploading new observations or to pause new identifications whilst seeking clarification and clarity. It’s quite something else to delete an entire account and everything associated with it.

I’m not sure what real benefits this might have. Sounds like the distilled essence of identification keys based on observation comments is one major aim. If so then it really does depend on how well such distillation can be performed. For rigorous work past performance of other similar systems isn’t encouraging.

isn’t this sort of a perfection-is-the-enemy-of-good situation?

when happens when a person writes something that is wrong? somebody sees the error, notifies the author, and the author issues a correction, apologizes, and moves on, right?

what’s the technical equivalent for the contemplated use case? you provide some sort of feedback mechanism to allow folks to report errors, and then try to fix the error by improving the training process in some way, and then push out new (hopefully corrected) information in as a short a cycle as possible. maybe you even add a mechanism to display a user-created correction note until the information is corrected. i’m not sure if there’s a great way to issue an apology and issue a formal correction for every such case. but is that necessary? and is anything about the contemplated use case so critical in the grand scheme of things that absolutely no errors can be tolerated ever?

in the early days of home computers, they would crash all the time, etc., but overall they were useful, and we improved them over time, right? i can’t think of the last time my computer or phone crashed. if we had just given up on them in those early days, i guess we wouldn’t be having this conversation at all because iNat wouldn’t exist, etc.

i sort of think that for folks who are concerned about misinformation, what we should be talking about is what is the right level of correctness to target before it’s useful enough. if it was wrong, say, 90% of the time, i think we could all agree that that would not be much success. maybe at that level, you would abandon the idea altogether and move on to other things. but then what is enough correctness to at least be interesting and maybe continue to develop? 50%? 60%? and what is enough success to put this into the hands of the masses?

maybe you also think about what kinds of errors are more important than others. for example, maybe occasional bad identification notes are okay, as long as they are never attributed to you. so instead attributing identification notes that could be potentially wrong to specific people, maybe the right way to approach this is to just point to observations, and let the conversations in the observations reveal who’s speaking.

The fact that my consistent opposition to generative AI throughout this discussion has somehow been interpreted and summarised by AI as being “relieved and appreciative” is VERY funny. Truly this technology is the future :P

Hey, at least we’re getting a sneak preview of what genAI applied to the rest of the website will be like: Falsified information directly attributed to the contributors.

Gotta say, there’s something downright dystopian about saying you don’t like an AI, then AI saying “look, everybody’s happy, even you! Here’s you saying how happy you are: [hallucinated optimism-bias algorithm disinformation]”

Thumbs down.

Please, at the very least, make this opt-in, not opt-out. People should be able to decide if they want their content to be used for this, especially given the strong reaction.

Seconded. This is actually worse. It feels like they have no idea what they’ve gotten into at all. How the hell do you mix up GenAI with anything else? Words have meanings! Facts matter! If it’s not GenAI why would anyone suggest that it was?!

Add that Google is, as is admitted in the post, swiping everything from everywhere, if they’re this ignorant about how it works, they’re also poorly positioned to prevent theft of their data.

I’m still deleting. It sucks, but if this grant is so small that if they don’t like the results they’ll just toss it out, they could also respond to this (pretty solidly negative) feedback by giving up on it right now. They’re not, they’re just spinning it differently.

Nope nope nope.