AI edited vs AI generated images — best practices?

Hi folks,

I am hoping to continue a recent discussion that began on a flag about assessing and addressing AI-generated vs AI-edited images. (I don’t think I can link the flag here per forum guidelines?) My understanding is that AI-generated images should be flagged as such, while AI-edited images should be DQAd using “Evidence accurately depicts organism or scene.”

Determining the line between AI-generated and AI-edited seems…difficult-to-impossible at times. For example, you can run an image through Gemini to see if the image has a SynthID embedded watermark, the presence of which indicates that the image was created or edited using Google AI. But because Gemini doesn’t differentiate between AI-edited vs generated, it can be difficult to determine how to proceed (i.e., flag vs DQA) when the organism is depicted in a relatively realistic manner.

Is there any sort of guidance available on this? Or failing that, any recommended best practices or common understanding re: how to handle these situations?

First, I believe you can link to flags here. It’s linking to observations or user accounts that is more problematic.

Second, there is no clear line between AI-edited and AI-generated, nor is there even a clear line between traditionally-edited and AI-edited these days. Many seemingly mundane photo-editing tasks like noise-reduction, dust-removal, contrast and color adjustment are done with AI tools these days. But perhaps what you really mean is “generative AI”, which should indeed be DQAed with “Evidence accurately depicts organism or scene". Unfortunately, there is no way to completely reliably determine when generative AI has been used. You just have to use your best judgement on a case-by-case basis I’m afraid. There are some tools that are supposed to help with this (like https://www.zerogpt.com/ai-image-detector) but none of them are foolproof and they often become outdated quickly.

The question isn’t “should things be marked as AI-generated” but “How should things be marked as AI-generated”? Should they be flagged, or should they be marked using the DQA?

Do staff or curators need to be notified, and take action, whenever someone uses AI to edit something? If so, then flagging is the best way to bring something to their attention.

Or is using AI to edit something discouraged, but not strictly prohibited? If so, then marking it via the DQA will make it casual, but the observer is free to go keep doing what they’ve been doing.

Sorry, please don’t post links to flags like this here. We don’t want people pointing to potentially negative behavior to any specific user and have it being discussed here, and linking to a flag does that.

Noting that this blog post clarifies that AI upscaling warrants a DQA flag, and I think there’s at least some level of debate as to whether that’s considered generative AI or not

Dang it, should have stuck with my initial instinct. Sorry! I deleted the link.

No worries!

I agree the line is a gray one, but the way I set it up is that the flag should be used for anything that basically didn’t start out as real photo of an organism. Whether that should be how it’s set up can be debated.

Thanks! Part of what I’m trying to understand is, in instances in which we know generative AI was used in some way (e.g., a SynthID is present), but we’re not sure whether it was used for image creation (thus warranting an AI flag) or limited to edits like generative fill (thus warranting a DQA vote)…should we lean towards DQAing or flagging?

DQA lets other users weigh in, while an AI flag hides the image from most users. It’s a judgement call. I’d factor in context (new user or not, severity of the edit, other flags on the user’s content, even rarity of the taxon and my own familiarity with it, etc.)

I agree that drawing a bright line here will be difficult. I think flagging is often good in cases where there is a pattern of behavior because it lets curators discuss and the user chime in as well if needed. If it’s a one-off then the DQA (if applicable) is fine. It’s also generally worth asking the user to clarify.

Thank you all — appreciate the guidance!

On some observations, I have seen the observer modify (unsure if using AI or just regular photo editing software, unsure how to check) the background to isolate the organism so that it appears to float on a stark white background.

I think this may be done because they find it aesthetically pleasing or to make it clear what the organism to be identified is or maybe even to make the CV work. I am not sure exactly why, but I do not think it is for any poor reason, as when I have seen it, it is not on all of a user’s observations, just on some.

Nonetheless, it makes it hard (at least for me, a less experienced identifier but trying!) to use clues about relative size or the plant an insect was on, etc. If Staff / Curators / other Vols could come up with a cut/paste to address that, it would be helpful, at least to me. :)

Thanks for keeping up with all these technological advances.

Specimen photos with a fully black or white background are pretty standard in biology since it often makes morphological details clearer. But for the reason you mentioned they should ideally include a scale bar (but yes, I’m quite guilty on not adding scale bars to mine, lol). While for most taxa these photos are usually made of a collected and preserved specimen, there are non-invasive ways to take them as well. Usually the method to take them is shoot against an already black or white background, but turn that fully black/white in post, but it can be done on some in-situ-photos as well.
Here’s one of the projects that collects observations like these: https://www.inaturalist.org/projects/inat-photo-ark

While I agree, I would tentatively put the line where data that wasn’t in the original file is invented by the program, for example, AI upscaling. Photos in which data has been removed (such as in the specimen-picture example) as well as images in which it has been compiled from multiple different photos (such as image stacking) should always be accepted, in my opinion, even if the program uses “AI”.

I have seen specimen photos (on the CICY website, for example) however none of the observations in the project you linked has scale bars that I could see, nor any information about their surroundings, such as host plants, plants they were found on, etc. Perhaps that ought to be part of a copy/paste?

In your post you say “it can be done on some in situ photos as well.” That sounds like what I am describing. Is this AI? I do not know. But the reason you give “it often makes morphological details clearer” has a tradeoff in that details are being lost, too, because these scale bars are not being added in, nor even any notes about size, information about surroundings, etc.

Sampling is prohibited here, so perhaps this also comes from that but I think I love nature perhaps more than biology, because actual specimen photos look flat and sterile to me, so fudged ones do not look any better. I have been skipping them as I see them and will likely just continue to do so.

That’s completely fair, all I’m saying is that since there are no details artificially added, these sort of observations should be allowed, even if the background is removed with AI.
While for my photos, I just touch them up by hand, I don’t see a big difference here, apart from manually doing it being neater at the cost of being very time consuming.

I agree with

myself, but here’s a real example that I see on iNat to illustrate why I don’t think this criterion is a “bright line”:

I ID primarily anoles in the U.S. - most of my IDs are for about five species (and then for some of the species commonly mistaken for them). I can often (but not always) ID fairly low quality pics to species because I know what species are in an area and what key features to look for. I’ve started to see low quality pics (far away, blurry, nothing too terrible) that look like they might be anoles, but are off - they often have stripes down the body when they should not. But they also don’t look much like other lizards that should be present in these locations, and, since the photos are low quality to begin with, it generally isn’t possible to be sure.

I’ve found that these photos all have one thing in common - they come from recent model Samsung Galaxy phones (generally 23 and up). In looking online, it’s clear that many of these phones have aggressive, on-board AI processing of images enabled by default (see the famous Moon-Gate for one example) that does some type of generative fill/enhancing. I don’t have a Galaxy to play with to see exactly how this might work, but based on this pattern, it seems very likely that these Galaxy phones are creating striped patterns on lizards that don’t have any in real life. The users posting these photos on iNat may not even be aware of this and are surely posting in good faith.

So how should these observations be handled? There’s no smoking gun that generative AI has been used, but it seems very likely. This issue likely affects many photos to greater or lesser degrees, as Galaxys are reasonably common phones, but the issue appears to be variable and stronger on lower quality pics. Conceivably, however, any pic taken by a Galaxy may have some generative AI content to some extent.

Personally, my approach has been IDing these to whatever level the ID is certain (often genus or “Lizards”) and ticking “good as can be”. I know the AI processing didn’t invent the lizard from whole cloth, but the actual details of the lizard don’t correspond to any species in the area. I probably would be justified in downvoting the “accurate info” DQA, but this makes an observation casual whereas a “good as can be” will at least lead to an RG genus level observation. I generally leave a comment explaining the issue, though not always. I guess maybe using the DQA would be appropriate and do better to futureproof that these images don’t get in the CV training dataset or something? But I am really not sure how to draw a clear line in this case.

Are these not a good use for downvoting the “Evidence accurately depicts organism or scene” DQA category?

It could be? My point is that this is going to be a sliding scale. If Galaxy phones are doing this by default, there are situations where they are doing it a little, a lot, and everything in between - the line of when the evidence is “accurate” or not will depend on the viewer’s perception/judgment call.

To my mind, whether or not the evidence is “accurate” also depends on the ID. In the case I wrote about, I am 100% sure the scene depicts a lizard (my ID) - it’s accurate to that level. It is not accurate to the level of species. If I encountered an observation like this where I wasn’t able to bump back the CID with my ID to a level I was sure was accurate, I would use the DQA.

The google pixel 10 is doing the same thing with photos by default. I had to uninstall it manually and then it limited the scope of the camera. It seems like they want to discourage disabling the Ai generation. When I was testing it, it added more than just a few lines here and there. It even added AI word salad to some things. I find this concerning.

(You can universally change your Google settings in your Google account re: AI.)