When the CV has confidently mis-learned a species

I was doing some IDing and was feeling quite pleased that there was a moth that I could ID with certainty: Banded Tussock Moth (Halysidota tessellaris). Then I saw a comment that as adults, the Banded Tussock Moth can’t be differentiated from the Sycamore Tussock Moth (Halysidota Harrisii) from photos. I did some investigation and found plenty of confirmation. In my area, both have been observed, as they can be easily told apart as larvae.
The problem is that the CV always identifies the adults as H.tessellaris, with 99%+ certainty.
After exchanging messages with the taxon curator, I’m now going through all the ones in my city IDed as H.tessellaris and re-identifying them at only genus level as Halysidota. Then I’m marking the DQA as “As good as it can be”. I’m also spreading the word amongst my local iNatting friends, asking them to help out by raising their spcies-level IDs to genus.

There’s discussion on Bug Guide here … https://www.bugguide.net/node/view/714227

I’ve heard of at least one other situation like this (with earthworms) where the CV is confident in IDing something to species level but experts say it can only be accurate to genus.

Does this happen a lot with the CV? How can we prevent it learning from observations that have been incorrectly ID’d? Is there any kind of ‘override’ that experts can put on certain taxa?

by correcting those observations, either with a different species ID or pushing species IDs back to a coarse level (whichever is appropriate)

Unfortunately the situation with this species will never improve. These two species are easily identifiable as caterpillars, so they will both remain in the CV training due to the thousands of larva photos, and they will continue to be “expected nearby” due to all the confirmed larval records. Once you get up far enough into Canada, H. harrisii does not occur, so adults up North can be confidently called H. tessellaris. That means the CV is trained on many correctly identified adults from that edge of the range. There is nowhere where H. harrisii occurs without H. tessellaris though, so the CV doesn’t have many correct harrisii adults to train on. The result is we have two common widespread species with thousands of larva photos in the CV training, but only a bunch of adults for one of the two species. So although the adults are not identifiable from photos over most of the range, no amount of ID-correcting in the world will make the CV stop suggesting tessellaris for all the adults. This is truly a case where even training the CV on 100% accurate photos will still result in rampant incorrect suggestions due to the way the model works.

Yes, within certain taxa this happens a lot; others, I gather, are largely spared. There’s a feature request you may wish to support: https://forum.inaturalist.org/t/force-computer-vision-to-back-off-on-the-specificity-of-suggested-ids-in-regions-with-cryptic-or-hard-to-identify-species/70021

This feature request may or may not become superfluous once some improvements to the CV model, scheduled to occur by the end of the year, are made:

We’re completing work on CV improvements that will give higher-confidence suggestions at taxonomic levels above species — a family or genus rather than a potentially incorrect species — for observations where species-level confidence is low.

I am not convinced to that the announced changes to the CV will actually address situations like this. The description suggests that the plan is mainly to increase the confidence threshold for suggesting species without major underlying changes to the way that it is trained or determines matches. In other words, I expect that it would be an improvement for situations where there are multiple similar species that are all included in the CV, and possibly for cases that it currently identifies as the wrong genus because the relevant species are not in the CV.

But as long as the training doesn’t incorporate disagreements (pushing the ID back to genus) or some metric such as % of observations not RG at species, I think the CV will likely continue to be confident about species where the subset of observations that it has learned primarily consists of observations identified as that species.

CV isn’t perfect and never will be. It regularly tells people that Chusquea scandens (Andean bamboo) is “visually similar” and “expected nearby” in Costa Rica, despite many of the “visually similar” species looking nothing like it and C. scandens being limited to South America. It gives me something to do on the weekend, at least.

The number of such cases of inseparable adult moth and butterfly species on which CV has overconfidently placed one name is fairly widespread…and frustrating. As with the case of the two Halysidota’s, there is often a geographic component to such dilemmas. Off the top of my head, three cases come to mind which are prominent here in Texas:

Burnsius communis/albezens (Common/White Checkered-Skipper)
Bulia deducta/similaris (Deduced/Similar Graphic)
Petrophila bifascialis/cappsi (Two-banded vs. Capps’ Petrophila)

In the latter case, a genus on which I’ve published ID articles, across 90% of the wide North American range of Two-banded, the ID is slam-dunk easy in reasonable photos. But in a narrow zone from south-central OK down through the Texas Hill Country, we also have the endemic Capps’ which at present can only be separated with a good view of the hindwings. So far, iNat’s CV hasn’t “learned” that latter constraint. For their part, the moths don’t always perch with the convience of iNat identifiers in mind. Yet CV has trained on so many Two-bandeds elsewhere that it will always suggest Two-banded for Texas examples even though they lack a view of the hindwings. In this instance, I just perform the (annoying) task of moving such observations back to genus level and leaving a brief note why.

yes.
thanks for making corrections!

The CV has literally mis-learned hundreds, if not thousands, of spider species, also for the abovementioned biogeographic limitations of some sister species. My main work here is stupidly re-assigning observations back to genus or species-group. The CV needs to be drastically improved for many arthropod species.

If you’re doing a lot of moth ID in southeastern Canada, here are a couple more to add to the “do not confirm” list, which the CV is overly confident about in this region:

Amphipoea americana cannot be distinguished from Amphipoea interoceanica without dissection, but CV calls them all americana. These can be placed at the species complex level.

Coleotechnites florae was described from lodgepole pine in Saskatchewan, and no one has managed to examine the type material recently to confirm what entity the name refers to, but it’s likely that all the “C. florae” records from the East are something else. I just ignore anything labeled “florae” pending a new publication, because the florae IDs will all need re-examined eventually and there’s no reason to think they’re correct other than “the rest of the internet uses that name for everything that looks like this”. CV calls all the black-and-white Coleotechnites “C. florae”.

Clepsis peritana is inseparable in photos from Clepsis penetralis, which occurs with it throughout southern-central Canada. Even the male genitalia of the two species are inseparable, but weirdly, the female genitalia are so completely different that the two species are probably in different species-groups. There are 1000+ RG observations of peritana in the northern USA and Canada, and 3000+ at needs ID, and sending them all back to genus seems like too much work. But as IDers we can at least refrain from suggesting/confirming peritana when the CV is confident about it.

The native Oegoconia novimundi and the introduced Oegoconia deauratella are only separable by dissection, and they occur together in eastern Canada and the NE USA. They can all be sent back to “as good as it can be” at genus. CV seems to suggest either species at random for any given photo. They’ll always be in the model, because there are regions where only one or the other occurs, allowing for positive ID based on range.

And one to not stress out about separating, because there are plans to lump them:

Euchlaena muzaria and Euchlaena obtusaria are almost certainly variants of a single species, and which name gets put on which photo appears to be random. The supposed genitalia differences are a spectrum rather than a binary, and it’s probably one variable thing. Both are in the CV model, and if you were to shuffle the iNat photos of both together, there’s no way anyone would separate them out the way they’re currently sorted.

People need to remember, and new users told, that the CV system is not and ID system, it’s a suggestion system. Trust it only as far as you yourself can trust your IDs. By all means use it to check the suggested species though.

The only way to get better CV suggestions is to make accurate non-CV based IDs and to correct incorrect IDs.

this is all good and useful information.

Does anyone know if there is a list somewhere of all the problematic species pairs & the recommended course of action for IDers to take (e.g revert to genus, revert to species complex, etc.)?

If there are really hundreds or even thousands of situations like this & maybe staff can be persuaded to change how the model works or tweak it in certain circumstances. Would be more persuasive if there was a comprehensive & well-curated list.

A couple more tricky ones, although I don’t know how badly the CV misbehaves for them:

Symmerista albifrons, canicosta, and leucitys: Adults require in-hand examination of the underside of the abdomen. Caterpillars look identical up to the final instar, when S. albifrons and canicosta develop more black and white dorsal lines than leucitys but can’t be distinguished from each other. Only S. leucitys can eat maple, but otherwise, you’re out of luck. Section Symmerista albifrons is the lowest ID that includes all three. Their ranges in eastern North America are not perfectly overlapping and sometimes you can decide between albifrons and canicosta caterpillars based on range, but southeastern Ontario has all three.

Desmia funeralis and maculalis: Adults can only be distinguished by a ventral view of the abdomen. Caterpillars are indistinguishable. There is a species complex for the two (Complex Desmia funeralis). Ranges pretty much overlapping.

On the butterfly front, the adults of Limenitis archippus and arthemis are extremely easy to distinguish, but the caterpillars are pretty much indistinguishable in photos. L. arthemis can feed on plants outside the willow family in addition to the willow family hosts it shares with L. archippus. So if a Limenitis caterpillar is feeding on a cherry or something, you can say it’s L. arthemis, but otherwise the genus is safest. At least on the website, it seems like the CV has gotten better about just suggesting the genus for the caterpillars, but I don’t know how universal that is

Edited to add range comments

I’m pretty sure the short answer is “no, there is no comprehensive list”, partially because for most of these it’s not a universal recommendation but an if-then tree that has to consider multiple factors that potentially include life stage, phenology, location, sex, etc.

Some groups of identifiers or identification projects may have agreed on some problem taxa and responses that are relevant to them. (The ones that I posted are pretty well agreed on among the frequent identifiers in the Caterpillars of North America project.) But even then I’m not sure how you would go about getting or documenting an “official consensus list of problematic groups” within a taxon x region combination.

The challenge with this is that all of these problematic pairs are location-dependent. All the ones listed on this thread are southeast-Canada/Northeast-USA specific. For example, in Florida, the Symmerista albifrons/canicosta/leucitys issue is easy, because only albifrons exists there; and Clepsis peritana ID isn’t an issue, because penetralis doesn’t occur there. On the other hand, in most of Canada, identifying Anania tertialis isn’t difficult, but over most of the eastern USA, it should never be taken to species without dissection because it occurs with the externally-identical Anania plectilis over a lot of states.

Walshia miscecolorella shouldn’t be taken to species over a large portion of its range due to several externally-identical species, but there are a few areas where miscecolorella is the only Walshia, and it’s safe to take them to species there. Datana drexelii is an easy ID in Canada and New England, but needs to be bumped back to complex when you’re within range of Datana major farther south. Apantesis anna is easy to ID unless you’re in a Canadian Zone peatland, where the hindwings need to be seen to separate it from Apantesis speciosa. Etc, etc… I could probably list 100 more cases like this, but you get the point. I don’t know how “reining in the CV” could be accomplished, since the times when we want the CV to hold back are dependent on location and/or life stage, and that information is either never (life stage) or only sometimes (location) considered when the CV makes its suggestions.

This issue is impossible to fix without a large change to the CV, it would probably take a lot of work figuring it out also as its quite complicated.

At this point as identifiers, i second paul_dennehy by just ignoring these problem taxa as trying to keep up with correcting will easily suck all your time up and become exhausting as it is never ending correcting. iNaturalist is just not the place for some taxa to be accurately identified by the current systems in place. Hopefully this will change one day.

No, it is not complicated. I would go as far and say that the CV should stop making species suggestions in many cases, even when locally it’s possible to ID the species. When species look identical to the human eye and scientists dissect them to ID them correctly, there shouldn’t be a CV trying to tell us differently without an extensive, standardized analysis that was fed into the CV or done with it. Think about it, the current approach is utter madness given that everybody can ID everything and upload non-standardized images. The final ID step should then be done by the community by hand. As a professional scientist who actually worked taxonomically with such difficult species groups or sister species, I definitely feel antagonized by staff here not improving the CV.

And no, ignoring is not a good approach. These data linger around in GBIF, are used by students or even scientists and often pop up as second layers in record schemes or even together with the actual records. And many of these species observations are just, plain and simple, wrong.

I agree. Especially as one of the things iNat is useful for is tracking the spread of species into new areas, it seems wrong to say “It can’t be species X because that’s never been seen here before.”

On the taxon page, the “Similar Species” tab can be very helpful. If the CV commonly mis-identifies a taxa that means there will be a lot of observations of the other taxa under the Similar Species.

That’s also what happens to taxa that aren’t in the CV yet - filed under another species. But once you know this happens, you can easily find the observations and correct the ID if you want.

I think of CV IDs less as final determinations and more as a first-pass sorting mechanism. Putting them under a look-alike species is a lot better than leaving them as State of Matter Life!

Maybe I’m misunderstanding, but I don’t think this is true, strictly speaking.

It would take a lot of work to undo the mistraining and people’s propensity to choose a species, but by going back over old observations and setting a species back to genus where appropriate, the model will improve over time.

I’m not saying anyone in particular is “responsible” to do this of course, but it could be tackled by a few dedicated identifiers. I’d help myself, but am not a specialist in moths.

Part of the problem is that people misunderstand CV/ML as a form of learning or intelligence, when it’s really guesses based on previous training–these guesses can be incredibly good when the training set is, but when it’s muddled, can be frustratingly incorrect.

One example I’m still in the process of working on is Quercus agrifolia, for literally YEARS the model couldn’t even pick it up most times, calling it Eucalyptus, Cork Oak, even Ficus when presented with whole-tree pictures.

I had thought that variability in branching and leaves was the issue and that there was some sort of machine learning or more sophisticated models at work that went beyond just comparing to earlier images, but no, it’s really that simple–so by adding my own photos of the species and identifying ~10% of the taxa (~60,000 total observations right now), the model has improved slowly, and even suggests Quercus agrifolia at least somewhere in the list most of the time.

Another example included Cynoglosseae, with most observations of that tribe in California being called ‘Popcornflowers’ (Plagiobothrys). By going back and setting incorrect IDs and blurry unidentifiable photos to tribe, the model started suggesting Cynoglosseae at the top instead of any particular genus or species, and then improved as the experts in the various genera started correctly identifying them to genus or species.

So really, any problems with the ML can be sorted in my experience, it’s just a matter of dedicated patience and effort to make the IDs and educate others. Insects are a particularly tough example though, with a lower ratio of dedicated identifiers to the sheer number of species that are out there from what I can see.

Maybe @thirty_legs could speak to their experience with earthworms as well, both frustrations and progress with improving the model through higher-level IDs and educating observers in comments with identifications?