Allow some non-leaf taxa to be added to the CV model

A similar problem has been brought up again and again and again on the forum: There are specious genera (sometimes with hundreds of species) where most of the species cannot be identified by photos, but one or two distinctive species can. Once one of those species gets added to the CV model, it drops the genus, and those other hundred or so species spend the rest of iNat’s history being misidentified by the CV (either as the distinctive species or a totally different genus). This problem has become so unmanageable that iNat identifiers have started purposefully avoiding IDing certain taxa to species in order to make sure that iNaturalist doesn’t drop the genus from the CV model. For other genera like Xysticus (293 species), Ophion (96 species), and Lyssomanes (94 species), it’s too late and we now have to spend countless hours fixing bad IDs for these.

I think the best solution would be to have some way to create exceptions for these situations and allow some taxa to be kept in the CV model even when they aren’t strictly leaves in the CV taxonomy. This could either be a manual process controlled by staff or curators or based on some programmatic threshold, e.g. taxa with less than 5% of their children covered by the CV model.

Although i am a very strong advocate for change. The CV Is a complicated system overall already. What ever request made needs to be quite thought out and needs strong support from the community. What would be the most work, but best for the long term would be for allowing the learning of higher taxa in general. Probably up to something like family. But there are concerns to even this. If one was to include all of the parent taxa of the current leaf taxa learned. That could easily balloon the number of taxa the CV is trained on to many times the size it is currently. This would create multiple challenges.

One thing is for sure though. iNaturalist strives its self in being community run. The community should have a way to interact and participate with the CV if it becomes a large problem for the community. Nobody benefits from an out of control taxon suggested by the CV that is constantly incorrect.

To add to the above. I strongly believe in setting up a system where taxa that become very problematic for Identifiers can actually be flagged and discussed being dropped from the CV.

A possible simple workaround is to change the algorithm not to learn the species until there are some threshold number of species eligible for training. Right now that threshold is 1. It should be a higher number. I would suggest 2 to start. That would leave some genera with only a single species from being identified to species. I think that is reasonable, because since we know it has only a single species, we can infer the species by the fact it was identified to genus. So this doesn’t seem like a major issue to me.

Not a bad idea. But unsure. It does seem that any idea will have costs and benefits. This would solve some issues, but create others. But in cases where this is a real problem. It will not work for a number of cases. Many large problem taxa have multiple IDable species. While this would solve the Ophion problem, it would not solve many more.

I think something more robust is needed. But it is difficult to come up with ideas that work when so many different situations and individual problems can arise.

Ideally this problem could be automated but if it can be, it would be somewhat complicated and take some workshopping to figure out the right decisions to make, comparing ratios of species and observation numbers. This is especially the case with incomplete taxa where the full taxonomy isn’t loaded into iNat yet (e.g. there’s no way for the CV to tell if all species in a plant or insect genus are included on iNat yet). I’m not sure if there’s any way to automate it that wouldn’t just cause significant problems with obvious species IDs.

I imagine a solution could be made where there’s some checkbox (which only curators could modify) which would identify a genus as one whose species should be excluded from the CV. Then all that would have to be done when training the CV is to download these data and exclude them from the training set.

While we’re on this route, a solution like this should also be set up to opt-in hybrids. There are a bunch of plant hybrids which should be included in the CV which the CV is identifying as the most similar non-hybrid, polluting the non-hybrid taxa with all the photos of the hybrid (Crocosmia x crocosmifolia is one that the CV consistently identifies as C. aurea, leading to much frustration for myself)

Wouldn’t it be better to have the checkbox identify a genus which should be treated as a leaf taxon even though it isn’t? That way you should theoretically keep the CV’s distinct species while also giving the option of a genus-level ID for other species.

Alternatively, I wonder whether the ‘as good as it can be’ checkbox could be used for this in some way - maybe if enough observations at genus level are marked ‘yes’, it can be effectively treated as a leaf taxon?

I am confused here: what do you mean by “drops the genus”? CV clearly knows genera for which it knows individual species, because it is never “pretty sure” beyond genus level.

The CV is no longer trained on the genus if it has learned any species in that genus – that is, the training set does not include photos of observations with a genus level ID or observations of other species within the genus that don’t yet meet the criteria for inclusion themselves. Material used in previous iterations of the training does not seem to be carried over (it does not “remember” out-dated information).

However, as I understand it, it is provided with some information about taxonomic hierarchies so that it can compare relatedness of the taxa in its training set when making suggestions. For the website and the old versions of the app, the top CV suggestion is deliberately always something higher than species. But when it makes these more general suggestions, it is not because it knows the genus per se, but because it knows a particular species and knows the species is in that genus.

If the CV only knows a few species in the genus, but these species are more-or-less typical representatives of the genus, this would not be a major issue – it would suggest the correct genus even for observations of species it has not learned. But when the only species that are identifiable are atypical representatives of the genus, it does not recognize typical members of the genus because they do not resemble the material it has been trained on.

An example. If the CV is unsure and the top suggestions are Glyptotendipes, Axarus festivus complex, Chironomus crassicaudatus, Chironomus decorus complex… it would likely suggest Chironomus Group becuase that is the common parent taxon of all four. It suggests the Genus Group, but it isnt trained on the Genus Group. It doesnt know Baeotendipes, Enfieldia, black Chironomus sp, etc.

Uh, that’s a bit of a shame. Even completely disregarding the problem outlined here, wouldn’t it simply be better to train it at different levels directly?

For taxa where the vast majority of observations are IDable to species, I imagine it would not be worthwhile to train it on higher taxa – it would be able to suggest higher levels simply through comparison and not need to be trained extra on them.

But where a meaningful percentage of observations cannot be ID’d more specifically, this means that there will be certain types of images that it will be unable to recognize. This includes not only groups of lookalike species, but also cases where certain phenological or life stages are more difficult to ID than others (for example, larvae or pupae, which sometimes have to be left at genus even when the adults can be distinguished; or burrows/nests without an occupant that could have been made by one of several species).

I want to note that although I am one of the people who has been suggesting for some time that we really need to include higher taxa in the CV training, I am in no way a programmer and I suspect there may be technical challenges to training it simultaneously on both parent and child taxa. I would be curious about feasibility and whether the reasons for including only leaf taxa are primarily about efficiency/computing power, or whether there are other barriers to implementation.

Similar feature request solved to focus discussion here:

https://forum.inaturalist.org/t/allow-for-genus-level-cv-training-sets-irrespective-of-species-level-participation/63938

Hi all, just adding to the conversation here since I think it is very important to fix this issue. One of the main problems that I don’t think anyone has mentioned so far is the geomodel, which works in close conjunction with the CV. I believe it is possible to resolve a large part of the issue without even touching the CV, but allowing for the geomodel to exist for higher rank taxa (up to family). Which should be easier assuming that adding more to the geomodel portion will not be as intensive for the system.

At the moment, when a species or genus within a family reaches 100 observations, the CV and geomodel for the higher taxa ranks are deleted. For example, some taxa I regularly ID:

Family Philosciidae with some species above 1000 observations

https://www.inaturalist.org/geo_model/48294/explain Geomodel does not exist

Family Eubelidae with species below 100 observations

https://www.inaturalist.org/geo_model/475299/explain Geomodel exists

So, areas such as Asia, Africa and Oceania which do not have 100 species level observations for any species in Philosciidae, but have over 1000 observations at family level, the taxa will never be suggested. But for Eubelidae with less than 400 observations in total, they will. This is because the vast majority of species level observations are concentrated in Europe and North America

I’m sure this is the same for many other families that are well studied in Europe and North America but lack research in other regions. Frankly, this current system is neglecting many of the areas that are already less privileged.

Still, the CV is able to recognise species within a family, and if an observation looks sufficiently similar to species that are currently known by the CV then it is occasionally able to suggest them. For example: https://uk.inaturalist.org/observations/339377138

Here the CV does not recognise this animal as any species that are present in this area according to the geomodel, so it reverts to only using the CV, and it successfully suggests many species that are in the same family as the pictured animal. BUT, it is unable to suggest that family because the geomodel has no species in that family which are in that region, and the family geomodel does not exist.

Another example: https://uk.inaturalist.org/observations/338052162

In this case the CV gives the best match it has for species known in that area which happens to be a frog, despite the fact it is a woodlouse, and if you click include suggestions not expected nearby it is clear the CV knows it is a woodlouse because all but one of the suggestions are woodlice.

If the geomodel existed for the family then this issue would be effectively solved, without having to adjust the CV in any way. Hopefully this makes sense and is actually a reasonably achievable solution, or at least a step in the right direction for the CV/Geomodel system.

The CV and Geomodel are linked. If you have a geomodel but its not in the CV. All you have is a map that has no influence over anything.

The CV forgets parent taxa as it is only trained on eligible leaf taxa. This is in my opinon its greatest flaw.

When Family Eubelidae gets a lower taxon learned, it will be in the same position as Family Philosciidae.

Indeed, perhaps I didn’t explain it thoroughly enough. I propose a system where the CV is able to suggest family rank IDs based on the species level suggestions it is already giving, but taking into account the family rank geomodels (that don’t currently exist), by looking at the parent taxa of those species and seeing which are present in that area. Then, in the observations I linked, Philosciidae and other woodlouse families would be suggested by the CV instead of the frog.

I know this doesn’t fix all of the issues but it I think it would be a step in the right direction and wouldn’t be as much trouble to set up.

I suppose this may have limitations where certain species within a group look different to the species currently known by the CV, which only could be solved by training the CV on the observations with the parent taxa ID and creating a corresponding geomodel. That would be an ideal solution. I only suggest this workaround because I believe it would be easier to implement.

I’m thinking of the earthworm Pontoscolex corethrurus here, which is most identifiable in its invasive range (global tropics) but harder to ID in Caribbean/Central/South America where it has endemic look-alike relatives. In that American range, its family Rhinodrilidae is also full of large, pigmented species that look nothing like this small unpigmented pink worm. No other Rhinodrilidae will ever reach CV eligibility because dissection is needed to identify. This results in the Geomodel/CV thinking the Rhinodrilidae is a single species of pink worm largely found in Asia instead of a diverse family of large, pigmented worms endemic to South America. This problem is basically repeated for every earthworm family that has any members in the CV, and I assume many more taxa as well, for which a change of what taxa the CV learns at (or using the CV at all) is probably the only fix.

Yeah I think an ideal solution would be creating a CV and geomodel for every taxon (perhaps up to family rank) that has over a certain number of observations at that level, learning from the observations at that level. Though I assume this hasn’t already been implemented due to the strain it will have on the system.

I’m not sure if anyone has ideas on how to optimise this?

There have been suggestions on how to optimise this, for example: