Well, technically, the genetic fallacy would apply only if I failed to believe what the original tweet claimed to be true by virtue of the fact that iNat and/or Google made the claim. I do believe what the tweet claims to be true - that this is going to happen. I just don’t like it. :-)
Guilt by association is not a fallacy per se, it’s a heuristic that can be applied appropriately or inappropriately. I don’t think iNat has become evil by virtue of working with Google. That would be guilt by association. I think iNat is making a mistake in partnering with Google in this way because Google has demonstrated its willingness to punch its users and partners in the face for profit. I think iNat is hanging out with kids who are likely to punch them in the face.
No way I’m gonna stop using iNat. :-) I just won’t use the feature.
But that’s enough out of me. No one should care what I believe to be the case about hallucinations, because no one has any reason to think I know what I’m talking about. ;-)
So rather than take a step back and look at the overwhelming majority of iNat users explaining their opposition and concerns to this, iNat has decided to double down on it. I feel like absolutely no questioning or research was done on this. It doesn’t take much looking to see what the average user of iNat is and how said users feel on gen AI. We’re not Facebook or Twitter users who suck up to it, we deeply oppose it.
The stated reasoning for using this AI is something humans have done for centuries: Identification guides. Why not have a section on species pages where we humans can add that ourselves? Human writing can be vetted and easily edited. Most people would rather no information than unconfirmed data that comes from gen AI.
It’s the same reason it’s dumb that iNat pushes heavily to get an observation down to the species level. Many observations simply have bad photos and therefore bad data to get there. If a human can’t ID it, AI definitely can’t either. And this comes back to the whole identification guides. What happens if a species only has single digit observations? Gonna make an assumption on that?
Like I said earlier, species guides exist already that date back centuries. I’m sure people would make guides on iNat if you added the page for it on species. That way it’d be done by people who are actually knowledgeable in the species or have proper data. I imagine it’d be a much easier thing to program than AI.
And finally, it’s frankly disappointing that this wasn’t brought up with the community or funders. It seems that once an organization reaches a certain size, transparency deceases and communication slows down. I was seriously hoping iNat would be different since it was built and run by the community for the longest time.
The more I think about this project, the more I’d like to see discussion of user incentives. Will subject matter experts still have the same desire to offer corrections and identifications, if their labor is no longer directly attributed to them, but is instead filtered through a process?
I say this as a professional insect taxonomist who assists with identifications at a moderate to slowish pace. Of course, this is all unpaid work. My employer doesn’t cover this. iNat doesn’t cover this. It’s just a hobby.
Some researchers will still happily contribute, I suppose, if their own research uses iNat data, and their labor helps clean the data. Not sure this is enough. Most people like some visibility or credit for their work, especially if it’s being done for free. What are the iNat incentives to keep the experts engaged? Because I gotta say, the prospect of posting knowledge that took me years to acquire into the void so that my words can get reprocessed and blended and presented stripped of any human element is a bit bleak.
I cannot find words strong enough to express how opposed I am to this idea. For what it’s worth (not much), I fully plan on wiping my observations and using other services if generative AI/ a LLM is implemented how the announcement suggests.
The $1.5 million iNaturalist is getting from Google won’t be enough to keep the lights on if their users leave en masse. It’s happened to other thriving communities and iNaturalist isn’t immune.
In hindsight, looking back at previous topics I had created expressing my concerns about generative AI or scraping of iNat observations, I get a sense that iNat staff and moderators were trying to shut down discussions as quickly as possible, locking my latest topic on how to prevent iNatters’ Youtube videos being used by third party AI, stating it wasn’t nature-related.
Google owns Youtube.
Call me deeply cynical but I think that staff have been trying to guide the community for years towards accepting generative AI.
Edit: The experience left me with a bitter taste in my mouth and less happy to participate in discussions on here as I deeply disagreed with that rationale, but I reasoned that oh well, they’re their own organisation, they can censor what they want. Now I cannot help but ponder if moderation decisions were or will be driven by acceptance of the grant overlords.
What about when people quote field guides in their comments? The author of those field guides didn’t give permission for their text to be compiled and used by the AI.
I’ve noticed quite a lot of instances where students copy text descriptions of organisms and put it in the notes or comments, but that description does not describe what is actually in the photo. They just take the CV suggestion and copy paste info about that species. I don’t think they even bother to read the text they are pasting in or they would notice the disparity. Seems like this could cause problems.
whatever the guidance system is, it effectively has 2 types of users it needs to interface with – those who need the guidance and those who provide the guidance.
the issue with approach 1 is that it’s not easily scalable. you need someone (or a group of people) familiar enough with a taxon, across different life stages, different seasonal forms, etc., but then they also have to understand how to input that information into the wiki platform. they need to know how to make the information useful enough for beginners and experts, for people speaking different languages, etc. ideally, they need to know how to get the best pictures relevant to a particular situation. that’s hard enough to do achieve a single taxon, but now you have to do it across thousands of taxa.
an AI-based approach is challenging for different reasons. it’s much harder to build that than a wiki-style system, but once it’s built, it addresses a lot of the issue with approach 1 because you don’t have to have the experts do anything different or in addition to what they’re already doing. the AI will just scrape the information they’re already providing elsewhere, and it can theoretically change the information it’s providing the person who needs guidance, too, to match their level of knowledge about the subject.
i don’t think your options 2 and 3 have to be mutually exclusive (in the long run), but i think if you’re just starting to develop something, it makes sense to start with 3 because it’s a relatively well-structured data set where the text generally should be relevant to the accompanying pictures. compared to iNaturalist identification notes, the language in scientific literature tends to be very technical, and it doesn’t always come with photos (often you get no images or illustrations instead), or it can be hard to figure out sometimes which image goes with which text. so i think it’s a less ideal starting for point for creating something to guide those who need the most guidance.
the other thing about an AI-based approach is that once you have something that “understands” how to identify something both visually and conceptually, it can eventually be used as a starting point for applications other than just guiding people how to identify things. as i noted earlier, it can be used to guide observers to get relevant features as they are observing. and then outside of the scope of iNaturalist, you can use that to improve automated monitoring of stuff for all sorts of real-world use cases, or you can use it to point out identifying marks that may have been missed in the literature, etc…
just looking at the example in the blog post, it looks like the goal is to create something very easy to use for beginners. to me, it looks very similar to the species comparisons from the All About Birds website (ex. https://www.allaboutbirds.org/guide/Painted_Bunting/species-compare). i think that’s a great way to begin, but i hope that at some point, it could be advanced enough to offer examples of observations that contain good examples to look at and also pull out links to guides and keys that folks have referenced, offering those as options for further study.
i fear that without offering sources, it’s going to be harder to find and close any misinformation loops that occur. it’s also potentially going to feed the narrative that this is just another AI appropriating human knowledge and putting it in a black box. to me, one of the key things that makes iNaturalist unique is the community. so i think it’s important that knowledge gained from the community can be tied back to the community as directly as possible. making that tie lets folks know that there are other people behind the knowledge, not just a machine.
as a former writer of survey questions, there seem to be too many conditions / variables in a single prompt, and choices offered here seem to have a bias to them. if you’re really hoping to gain some new knowledge from the poll, i think it would be helpful to rethink the way this poll is presented.
and by the way, i’m not trying to pick on you specifically, but i do remember a case involving us and an AI that sort of parallels the issues being discussed here. i thought it was interesting there that your AI improved my code in some ways (linting, standardization, etc.) but might have discarded a line of code that addressed a particular unusual use case. the fact that it didn’t refer you to the original source of the code means that you missed out on some of the usage notes, too. i personally didn’t mind that i wasn’t credited in this case, but if i had produced the code as part of my livelihood or my professional body of work, i can see how folks might be annoyed that their knowledge is used and adapted without at least a citation.
Copy/pasting my comment in the blog post with some extra formatting:
While I understand the reason, I think the level of the knee-jerk reactions in the comments is not proper of a platform like iNaturalist.
I too am deeply tired of LLMs being shoved down our throats.
I am, also, a person in the exact field where I’m being literally asked to work towards replacing my own position with an LLM. I don’t like LLMs. I also thought “Ah sh*t” when I first saw the forum post’s title.
With that said, I ask you all to pause for a second and consider this before posting:
Lots of comments are around “how about a wiki made by humans”, about letting the platform’s users do the heavy lifting. All taxa have an “about” section that links directly to the taxa’s wikipedia article, alongside links to other relevant platforms with info on the taxa. If no wiki article exists, there is a link to create one. Did you edit the wiki articles for every single taxa you offered relevant comments on before to ensure the info is there? Is there any instance of an “This is XYZ because of these reasons” comment that you left without ensuring that same data was added to wikipedia or other relevant sources of ID info for the public? Do you think everyone else also ensures the About tab has all those details every time they comment? Do you realise this project aims to do that work that you, me and millions of others didn’t do?
What would your reaction be if Computer Vision didn’t exist yet and it was announced today? Do you think Computer Vision should be removed entirely, since it also tends to give out wrong IDs? Should Seek be binned as well?
Did you take all possible actions on your side to ensure iNat doesn’t have to resort to taking a $1.5M grant from a megacorp? Did you chose the copyright option that allows everything to be uploaded to GBIF? Do you donate regularly? Do you campaign for others to choose CC0 and donate?
…If you did allow GBIF to use your data, do you realise that that action has already enabled Google to use it without compensation, permission or attribution as well?
When you say that this will erode the community, have you considered that this use would appear a step before the community gets involved, in the same step that CV appears? Do you prefer for non-experts to have less information at that step, in case some of that information was taken from an user that was wrong before, because you think that iNat users (the same ones that were scraped for the data) should be the ones giving out this information? Isn’t that circular reasoning? How do you know that the future IDer will be “more right” than the previous IDer whose info was scraped?
If you think threatening with leaving if this is implemented… What does this achieve? Does removing your contributions to science improve the science? Is “not making my hard work available to Google” (it already was) a good reason for making that hard work unavailable to researchers? Is making LESS good data available a good way reducing the enshittification of data?
THIS! All of this. This is why having AI try to write guides or help ID is terrible. Remember when Google AI suggested using Elmer’s glue for pizza? It said so confidently. A human wouldn’t do that. We say “I’m not sure, but” or something like that.
For me, it’s simple: I am yet to see a single instance of Generative AI being useful. I already don’t particularly like the Computer Vision because of how much confidence people put in it, and I can’t imagine a GenAI would be anything but worse.
Even excluding the enormous ethical issues of GenAI, the problem is that it’s just not accurate. It just spits out something that sounds like an answer, and even if that answer does turn out to be correct there’s no immediate way to distinguish between truth and nonsense, and equally importantly there’s no accountability. Even if the error rate was 1% (which it most certainly is not), the problem is not that such errors exist, it’s that they’re presented as definitive fact by a machine that doesn’t even understand what a fact is.
I’m certain that I myself make countless identification mistakes on iNat, but if someone questions me about it we can have a discussion and then we both leave a little bit more informed. Sharing information is a huge part of my iNat experience and it’s what made the place so great for me. It’s been a great place to escape the chaos of the rest of the internet where GenAI has been relentlessly shoved down our throats.
I agree with other user’s suggestions that a Wiki or something would be a much more useful and informative way to approach this problem. But if we absolutely must use AI, then a much better use of it would be to identify and link to observations or comments that have helpful ID tips. That way human users don’t have to trawl through thousands of observations to find useful information, but we also don’t run into all those problems I mentioned above.
For what it’s worth, if this does get implemented then I will have to seriously consider my future on the platform. iNat has been a fantastic resource but this just doesn’t sit right with me at the moment, and I don’t want any of my identification comments condensed into an 80% correct summary when I’ve already created countless easily accessible summaries.
I think context is important, here. The standards and requirements for Wikipedia are surely different than what any standards on an iNat Species Wiki would be, simply because their purposes and target audiences would differ.
When was the last time you saw diagnostic illustrations for a species on a Wikipedia article? I’m not saying it’s a bad idea to, if you’re able to, contribute to an “identification” section of any given species on Wikipedia. Just that context matters, target audiences matter.
I ultimately agree that we’d all be wise to keep cool heads about this. I’m pretty tired of LLMs everywhere, and I’m still not super thrilled about this decision. But being reactionary isn’t productive, so I’m trying to be patient, here.
Thanks for the link @natev. That’s an interesting read. I’ll note that the paper is based on comparing
Human-to-LLM ratios in terms of the energy consumption, carbon emission, water consumption, and economic cost for writing one 500-word page of content.
The authors also note that
A pre-trained LLM, like the popular LLaMA model families, is used by many users and downstream tasks including customized fine-tuning to suit a variety of applications. As a result, the total computational demand of LLM inference can exceed that of training by far.
And
To estimate the environmental footprints of humans, we assumed that each page of content has 500 words and that the average human writing speed is 300 words/hour, resulting in 1.67 hours for a human to write one page of content.
The proposed iNat scenario differs in several respects:
It’s not clear that the stated goal of iNat’s project with Google.org—to “tell people not just which frog it is but why it’s that frog”—can be accomplished with a pre-trained LLM. It may require a custom-trained vision-language model for which the training resource costs would be specific to iNat.
The generative AI model ultimately used for this application may require more or less effort to “explain” why a particular image best fits a particular identification than Llama-3-70B requires to generate 500 words of text.
On average, a human identifier probably requires a whole lot less effort to provide an ID, even with some explanation of distinguishing characters, than to generate 500 words of text. I know that I do.
It’s fair to say that the widespread reporting on the high resource demands of AI model training does not accurately predict the likely impact of this particular project. However, I don’t think we can confidently say what the environmental cost would be for an effective system to offer “AI-generated identification tips”.
I do, however, think it’s worth pursuing the project to find out (although I’m concerned that the predictive value of iNat comments may be a lot lower than if actual taxonomic papers were used). I would also love to see some analysis of the environmental impact in the later stages of the project, once it’s clear what the production process will be. One thing that moderates my concern is the knowledge that iNat’s budget will likely constrain the amount of resources that this functionality can be allowed to consume. I’d be surprised if iNat staff would find the identification tip process worth implementing if it has an annual cost anywhere near the current cost for CV training.
On a different tack, I think the distinction spidercat makes here is important…
As a long-time contributor to both projects, I completely agree that they have very different standards and work processes. Adding content to Wikipedia rightly requires a process of citing sources that demands quite a high level of effort. I believe some type of user-friendly, crowd-sourced, multi-access key could be created in iNat that would still be reliably sourced, but that would allow data to be contributed in smaller quanta with a lower time investment to achieve a given result.
Summarising a bit what I replied to somebody else in the blog post comments - I understand Wikipedia is not the same as what people are proposing to use instead of an LLM. What my point aimed at is:
Wikipedia is what is there right now. Yet, most people won’t even fill the most basic information when an article is missing or just a snippet. I see this even for charismatic birds. What makes us think it would be different for a whole new internal wiki started as a blank slate? Why is starting this new internal wiki with processed existing input inherently worse than the blank slate option? Is no info better than potentially wrong info scraped from community data, when we are talking about the INITIAL ID proccess, way before community IDs come into place?
Comments correcting other people’s wrong IDs, or responding calls for help will naturally be much more prolific than a passive request to fill in data for a new tool. Giving users the ability to correct wrong outputs would also harness our natural affinity for, well, correcting others.
Not exactly to your point, but let me add: Inat is a data aggregation tool. Inat’s culture has always been to allow everyone access to your freely given data. Even with my personal disdain towards LLMs, the mere idea that aggregating user data to build a tool (any tool) is against the user’s wishes is almost silly. From the moment you freely give access to your data, you are consenting for that data to be used within the limits of whichever licence you choose - by anybody, including evil megacorps and Anish Kapoor.
I will say that Wikipedia moderators are often complete buttholes when it comes to allowing you to create an article for a species. Often times it’s “not relevant” or “not important enough” and they just delete it. Also some people find using Wikipedia or similar sites hard, editing and making new pages isn’t casually easy. I used to be an editor for WikiFur and it took me weeks to find my bearings.
Here on iNat, that’s not the case at all. There’s nothing stopping you from straight up editing species lists and whatnot. And most things here are easy to do, usually a matter or one or two clicks.
Of course, there needs to be some restrictions. Like needing to have X amount of IDs or be a member for X length of time to prevent new members from vandalizing stuff. But that’s something Wikipedia already does and could be implemented here.
AI should not be entrusted to give detailed explanations on how to ID species. 80% of Chironomids haven’t even been IDed on INaturalist. How on earth can it be expected to generate anything close to accurate?
I fear this will just push many large identifiers away making so many of the issues of data quality, CV accuracy, this generative ai thing worse.