So the person pointing out the fraud may themselves be a fraud (or is it just me that creating your own journal and publishing in it isn’t at least borderline fraudulent?). The email on the “journal” is kka6@cornell.edu, which is tied to Kriston K. Anderson, Cornell alumni. This person evidently created a “journal” in which they publish.
The Journal does seem to be self published, with most articles by Kris Anderson. I checked out several volumes, but he still raises some worthwhile points.
I’d like for us to keep this discussion focused on the broader challenge of AI-generated images and iNat/science. While I agree discussion about the article and public information about the author are probably fair game, I think it’s more productive to discuss the bigger picture than one this one example put forth in the paper. Thanks.
Another related (recent) paper, about “fake” images, eBird, and detection patterns
Assistive AI and Data Manipulation in Citizen Science by Xinming Du :: SSRN
Fully agree, and to add to the point about the many details making Inat observations relatively hard to fake plausibly, I’d say that observations with multiple photos showing different angles or distances would also be difficult to create with AI - at the moment.
I have seen numerous cases where blatantly stolen photos have gotten IDs by multiple people before someone noticed that something doesn’t seem quite right and investigated. So I am not entirely convinced that faked and stolen images necessarily receive more scrutiny on iNat than elsewhere – at least, not the sort of scrutiny aimed at assessing the authenticity of the image rather than the ID.
How would this be enforced – i.e., how would you prevent people from editing images before uploading? Would this include things like cropping, adjusting color balance, etc.? How would iNat handle the additional burden on its infrastructure that would result from people using it for image editing?
Um, the iNaturalist website already has a tab that displays EXIF data, assuming that the user did not remove it before upload.
Yeah sure, not saying the opposite, just that maybe we should regard it more like “someone said something in a blog” that “a scientific paper showed that there is a problem”
Out of curiosity, what do you mean by ‘potential grounds for suspension’? Is there a strike system for users who’ve uploaded AI generated content?
To zoom out a bit: as others have said, people have been posting stolen photos to iNat forever. While people posting artificially generated images is a relatively new, and currently rare, wrinkle, I think it’s pretty similar to the former issue. In both cases, the evidence presented is not of the claimed encounter with the specific organism depicted in the image.
We then get to the motivation for posting such an image. In many cases it’s not malicious, it’s because the person couldn’t get a photo of the organism and they think they should post a photo of what they think they saw, even if it’s not theirs. The new observation workflow and the data quality grade system incentivize this a bit, so it’s understandable, I think. One of the Community Guidelines is “Assume others mean well” so unless it’s clear the person is doing this maliciously, and in my experience that’s fairly uncommon, they should first be warned that if it continues their account will be suspended. So yes, usually there’s one “strike” before an account is suspended.
Furthermore, I think it’s often not clear that the person posting the image knows it’s artificially generated. They may have just taken it off the internet without knowing its provenance. On iNat we ask everyone to only take their ID as far as the evidence can support, and I think we should extend that thought process to these situations as well. It seems like there’s an underlying assumption in this discussion that someone posting an artificially generated image is inherently more insidious than someone posting someone else’s photo but I don’t think that’s necessarily the case.
I agree that artificially generated images definitely present new challenges, like you won’t be able to use a reverse image search to find the image elsewhere, so there’s more uncertainty involved. But in my experience reverse image searches don’t always work for images that I’m 99% sure were taken from somewhere else, so even non-AI images present challenges.
Finally (sorry for the long post), let’s keep in mind that iNat already has a suite of tools that allow anyone to rate the data quality of any observation, and/or flag evidence. It’s always been possible for someone to post false/erroneous data (whether intentionally or not) and it’s always been possible for the community to weigh on that data.
Thanks @GeolTel. That article is much more rigorous. I don’t use eBird, so I was quite surprised at the apparent extent of the fake observation issues Xinming Du uncovers in Brazil:
rare species records in ecologically implausible areas increase by about 22 percent post-AI, while the share of these records flagged for expert review falls by 9.7 percentage points
Most of the paper is focused on statistical analysis of data anomalies post-2021, and I was skeptical whether these anomalies truly reflected AI use. It seemed hard to believe that there were a large number of Brazilians deliberately faking bird sightings. The paper does give an example that provides a little more context.
The Harpy Eagle (Harpia harpyja), one of the world’s largest raptors, is strictly dependent
on continuous lowland rainforest canopy. In the pre-AI period of 2017–2020, there were zero
eBird records of this species in municipalities with less than 5% forest cover. After the AI
boom, multiple new “first appearances” emerged in deforested municipalities in Maranhão
and Piauí, areas with less than 2% forest cover and no contiguous habitat. These records
survived eBird’s automated filters and were never flagged for expert review. This is not an
isolated case. Novel rare-species and low-forest municipality pairs increased from 262 per
year in the pre-AI period, 1,047 over four years from 2017 to 2020, to 315 per year post-AI,
1,574 over five years from 2021 to 2025, a 20 percent annual increase.
It’s still unclear to me whether the paper was able to separate submission of AI-generated photos from stolen photos, but the supposed appearance of rare, forest-dependent species in urban areas clearly indicates these records are largely fake.
Fortunately, we don’t seem to get anywhere near as many fake observations on iNat. As @mycographer says, the rate of fakery seems likely to be dependent on the incentives available versus the cost/effort to create the fake. Now that the cost has fallen near zero, it seems important to minimize the incentives (as well as improving detection).
Gamification within the platform seems like one possible incentive to fakery. Fortunately, the extent of gamification on iNat is pretty minimal. While there are leaderboards for observers and identifiers of every taxon, the sheer number of them means that there’s no one big competition that would bring recognition to a particular observer who saw some species the most times (for example).
Bioblitz-type events perhaps do create an incentive for fakery, as they typically give recognition to observers with the most observations or species. The biggest of these, City Nature Challenge, had clear problems with the quality of observations submitted over the past several years. So far, AI images haven’t been a major issue for CNC—instead it’s been the submission of stolen images, plus the huge number of cultivated/captive organisms. But the addition of AI fakes further emphasizes the need for project admins to take care with education and to closely monitor submissions. Here is the iNat Bioblitz Guide.
The other big source of problem observations is “duress” users, typically students who are told they need to use iNat for a project and given a goal to make a certain quota of observations during a particular period. The iNaturalist Educator’s Guide has great advice on how to run a class project in a way that encourages good observations and helps students learn important skills.
The danger in this (as it’s often in any fraudulent act) is what can the fraudster gain from it. The best thing about iNaturalist is that it brings no economic incentive to fraudsters. Fame and ego have traditionally been a common incentive but also quite limited in reach. In birding I’ve read quite few old cases, for example.
As a society, we have new tools that can be used for fraud so we’ll need to adapt and consider them, but it’s nothing fundamentally different than before. It does make manipulating at picture so much easier, but it still doesn’t benefit anyone from it more than before.
Another topic, but this is why I think iNaturalist has ben wise to steer away from gamification.
Maybe I’m missing something, but this paper’s conclusions seem quite dubious to me. Its results seem to mostly come down to “the rate of extremely unlikely observations went up in 2021, and that’s also the year AI took off in Brazil”. Sure, this could be because AI made it much to fake photos, but I’m not seeing anything in the paper that even suggests this. I can think of two explanations for this trend that seem much more plausible to me than “AI generated photos have caused a huge uptick in false reports”:
- eBird could have had an unrelated uptick in casual users in Brazil for any of a litany of unrelated reasons, especially given the post-pandemic timing
- AI is involved, but primarily in casual users using tools like computer vision and large language models leading to incorrect conclusions about what they saw.
Anecdotally the vast majority of extremely dubious eBird observations I see seem to come from well-meaning but confused novices. It does not seem like AI generated photos are a driving factor there, because in fact the vast majority of these either do not have photos, or have photos that are either unidentifiable or depict another species.
The situation in Brazil may be different, but nothing in this paper seems to suggest this is the case. In fact the paper goes out of its way to describe how many of these dubious observations slip through the filter. Why go through the trouble of faking a photo if your observation didn’t even get flagged? If your whole argument is that the driving force here is fake, AI generated photos, shouldn’t you look into how many of these dubious observations have photos attached, and how many of those photos show signs of being AI generated?
As someone who is more familiar with how eBird works, I’m pretty skeptical of that paper’s claims. Most eBird observations do not require review or evidence, so I highly doubt the AI boom has had a significant impact on incorrect eBird observations.
Theoretically yes, they should have something like this in the info of the photo:
| Make |
|---|
samsung
| Model |
|---|
SM-A307GN
| Software |
|---|
A307GNDXS6CWH1
But I am not sure how easy it is to fabricate
Let’s not forget Piltdown Man and the Fiji Mermaid!
Pragmatically it’s trivial to fabricate/falsify EXIF - there’s a tool that can be used to easily copy data from one image to another among other manipulations.
I regularly edit the EXIF data on my images for the use case of overriding my phone’s geolocation data with a GPS track gathered by my smart watch. EXIF just isn’t cryptographically secure, it’s just data that lots of legitimate/helpful workflows write and manipulate.
I don’t think those bird records are all fake.
My experience over 20 years and photo records show that bird species appear and disappear depending which state has drought or dust storms at any given time.
The math is definitely incorrect. From 1,047 over 5 years to 1,574 over the next five years is only just over 5% annual increase.
A more likely explanation, and something I see commonly (as an eBird reviewer in Brazil!) is that observers map their checklists imprecisely. For example, visiting birders who start a checklist at their hotel and let it run the entire day as they drive out to different forest sites. That’s one of the big advantages of iNaturalist over eBird: that it is possible, at least in principle, to map observations to within a few metres of where the organism actually was. eBird data can commonly be mapped kilometres or even tens of kilometres from where the birds were actually seen.
I remember recently being shocked to see a checklist at the hotspot for the city of Teresopolis reach a triple digit species count, only to open it and realize the list represented hiking through multiple national/state parks which had their own hotspots! Checklist was also set to complete with a duration of “1 minute”, ugh…
I hope you flag those checklists when you see them! They’re the worst.
