Creating a database for species information

Hey everyone, I wanted to see if anyone has already got a solution to a problem I’ve found myself in with identifications.

I’m trying to ID Philodendron (a genus in the Araceae) to species level, but there are over 600 accepted species. So I’ve started using a collection of notes (in Obsidian) where I take notes on species characteristics, differences and links to useful websites. It’s great that I can link species to each other and create a taxonomic network. However, I feel like there is a lot of information I cannot quickly access this way. For one, geographical information is hard to access. Let’s say I encounter an interesting Philodendron on iNat in Costa Rica, I’d like to be able to just filter by “species occurring in Costa Rica” to see which species are actually present. Also, because there are so so many species, I cannot put all the important taxonomic characters into a short summary. There are just too many characters that are important in different species contexts (e.g. species A and B differ by one thing, but species B and C differ by something else and so on).

I was thinking of creating a database, because that would allow me to more easily filter the species. (I’m comfortable writing R code, but I guess this would be overkill for this? But I’m totally unfamiliar with actual database stuff like SQL.) If I’d get a good solution for this running, I’d also like to have this information available online so other identifiers can profit from it as well.

So far, the note-taking has helped me already in being able to get a better overview and to remember many more species. But I’m sure other avid identifiers have found better solutions to these kinds of problems. So, how do you tackle this?

Here is a screenshot of my current note-taking system in Obsidian:

There are probably people on here who can give you much more sophisticated advice than I can. But no one has answered so far, so I’ll jump in. I suspect the first step is to reduce your descriptive text to a series of choices that can go in a spreadsheet.

Recorded from Venezuela yes/ no

Leaf mid-rib hairy / hairless

That kind of thing. That would allow you to pick a character which you can see in an observation and select all the species which have that character. Then pick another character and reduce the first subset to a smaller subset, and so on.

I think this is Fantastic! But part of me thinks why are there so many species to begin with? and don’t they all hybridize like crazy (This is the Arum family after all, extremely difficult taxonomically)? Perhaps this is why they are so tricky?

What you are trying to do is create a key to identifying each Philodendron to species. Likely such a key may already exist. If we know how many species exist, even approximately, then we know what features separate one species from another. Finding such a key is often a challenge to find.

If the Araceae family is complicated, it might be helpful to find how the genera are identified. There are likely lots of Needs ID observations in Araceae waiting to be moved to genus. Looking through these can be a great learning experience.

Thanks for your replies so far :)

@jhbratton This would probably be a good idea, yes. I’ve had already begun doing that but it is a massive task and I thought, maybe I ask around here first before I do a lot of work for nothing (i.e . if I’d want to change my methodology later on).

@norwichtim I’m familiar with aroids, yes (I’ve already done over 22000 IDs of them). And no, with new species being described constantly and most of them restricted to only very specific locations, there isn’t a unifying key to them. Instead, there are lots of out-of-date keys for specific locations…

@professor_porcupine there certainly are hybrids, but so far I’ve only encountered very few (like 1-2 hybrid species). It is indeed a very diverse group of plants, but that’s also part of the fun :)

Just to clarify though, my question isn’t really about aroids or plants. It is if there are better ways to store and sort taxonomic data for better identification, regardless of the specific taxa.

I hope so, but if there isn’t I hope you make one! It’s gonna be a lot of work & you might have to get help from AI to do the repetive work (of course fact check it all afterwards).

You might take inspiration form this site : https://nwwildflowers.com/flora/

I hope to create something similar but for any species compared with any other species, side by side like Leaf, fruit/seed, flower (upclose & inflorescence), plant growth habit, ect + all other applicable traits. For now the best is a Wiki.

Do you plan to extend this concept with all plant species? Something like Plants of The World Online but with ID Info & keys for every genus/species? Might need to take a HUGE collaborative effort with the whole INaturalist community + other communities. Phylogeny & morphology also need to be taken into acount for species ID (Phylogeny help frame limits of a species - even if species boundaires are fuzzy the phylogeny can help show that).
https://powo.science.kew.org/taxon/urn:lsid:ipni.org:names:30000199-2

Damn… bro you are goated! I think you are uniquely equipped to handle this task. Remind me to ask you about Arum Family plant ID. Since you been studying it, how diverse are the flowers & fruits in the family? You got Monstera, Duckweed, Taro, Jack-in-the-pulpit & so many diverse species it’s crazy!
I had trouble sorting the genera & tribes of some of this huge families Subfamilies

It seems the Tribes of Aroideae are kind of polyphyletic & hard to make a tribe for.

I’m another person with the same problem but no great solution at this time. I kinda feel we might need to develop some kind of community-led wiki taxonomic database. Anyhow, what I’ve done so far has mostly been to use spreadsheets. Here’s one I have for the genus Sisyrinchium:

There are about 240 rows (species/subspecies) and about 350 columns (countries and first-tier subnational divisions). But there are only maybe 12 columns for characters. It’s been pretty helpful so far (6,500 IDs of the genus), but I really notice the limitations:

  • I can narrow down the potential taxa pretty well, but I often have to go back to the primary sources to get to an exact ID.
  • I’m tracking very few characters and I’m generally limited to filtering for values rather than using numeric ranges, for example.
  • None of the character or geographic statements has a citation so nothing tells me why I think S. tinctorium has pendent, pyriform capsules or occurs between 1300 m and 4500 m elevation.
  • It’s entirely private. No one else can use this info 'cos it’s on my computer.

I’ve looked at key building software/sites a couple of times, but never got to the point of generating anything meaningful. IIRC, there was a U.K. developed free offering that takes spreadsheets as input that looked promising.

A few times I’ve resorted to writing my own key in Word, and while it’s a bunch of effort it’s pretty satisfying to have a usable document that looks professional.

But it’s not a database of character traits and distribution, and it’s not a multi-factor key, and I feel the combination of those is the holy grail.

Great to see your interest but why attempt to reinvent the wheel. We see it all the time, people do a lot of work, very useful website, then they move on and it is all lost and for nothing. This is exactly what WFO aims to do: WFO Araceae. Contact them, see how you can get involved and how they can help you.

Are you opposed to storing the 12 columns as observation fields and slowly “and’ing” them together? Once you and’d as many together as you’d like, you could see them as pins on a map.

Maybe what you need is more complex than that tho?

You’d be slowly filtering observation images with each and’d field… once you got to the end the 12 fields you’d have a grid of obs matching the choices from the 12 obs fields making up the key (and could also see map with pins of where the obs were made that match the 12 selections).

But maybe I’m interpreting the need completely wrong.

I think this approach is very good. Using tags for this instead of a spreadsheet may even make this very easy.

If you have a document of any type with a description of a species, you can just apply all the required tags. Then filter your search for specific tags and with each additional tag you filter by the number of species-documents should decrease until you either are left with only one or at least so few, that you can look through them all quite quickly.

It has the advantage that it’s not dichotomous, but at the same time retains usability since the sorting basically happens automatically, once you set the system up

Hard to use a common key across multiple users is the only downside. Obs fields with drop-downs would make it consistent across users.

Unless the custom app was doing the “create” and limited the tags to the packet used within the key.

I didn’t mean tags on iNat, but rather on a document/files app. But yeah, same downside still applies

Hey everyone, thanks for all of your replies!

@Davidia Somehow the site you linked is blocked for me. But yes, I know that there have been several attempts made to make taxonomy more accessible in Araceae, CATE Araceae being another example (now defunct as far as I can tell). And most species information I draw on is from this source via POWO. However, what these sites have done so far is still one step removed from what I’m planning to do. It is valuable to have taxonomic information for all aroids, but you still have to browse through endless complex and vague taxonomic texts juggling species information of up to 10-20 possible taxa in your head for any given location. That’s why I would like to have a database that I’m able to filter by location and characters.

@rupertclayton Now that you mention the existence of software designed for this, I had a look around. One software to create dichotomous keys would be DichoKey using the monographaR package in R. At a glance this worked quite well, however I don’t really see how you could embed location data into this. Creating a standard key with this is much easier than writing your own key in Word (as long as you know your way around R).

I also found Xper3, where you can create, edit, upload and browse existing taxonomic keys. It doesn’t really have much content online yet. But the editor and the interactiveness seem to work really great. You get a list of all created “descriptors” and can select which ones to use and what value they should have. This way the list of possible taxa shrinks for each descriptor value you select. One could add e.g. country as one descriptor and have all possible countries as values. Or in aroids I’d probably have something like “Pacific”, “Atlantic”, “Amazonian”, etc. So this key isn’t dichotomous, which is a great advantage like @eyekosaeder pointed out. Maybe I’ll do an example key for 10 taxa later.

I also have struggled with this and wish I could work on a database for Chironomidae species sources, information, etc. It is just too much literature. Where do you find info on X genus ecology? I’m not sure, its probably buried somewhere. It would be great to have a place where a species could have links / references to wherever information about it is published.

@letiziaw But you can already do that on POWO (where the data is available), search “Araceae” and “Location:Tanzania” and “Flower:white”:

https://powo.science.kew.org/results?q=Araceae%2CLocation%3ATanzania%2CFlower%3Awhite

Though I agree, the data needs to be more structured to make it more accurate.

Hi @davidia I wasn’t aware that POWO stored any data on taxonomic characters. Is there a list of which characters and states can be used? Do we know where POWO’s data comes from and how it gets updated?

Unless POWO is taking on a massive effort to code character states and locations based on published data, I’m doubtful that POWO would be the best place to search characters of this sort. It seems much more likely that a database that allows individuals to add data themselves is going to meet the type of real-world requirements that @letiziaw and I have in mind. For example this POWO query yields no examples of Sisyrinchium species with yellow flowers, whereas my improvised spreadsheet approach shows me that 134 of 242 taxa have yellow flowers.

I wasn’t aware of this either, it’s definitely pretty handy to have! Like @rupertclayton said, it doesn’t really solve the problem though. For neotropical aroids, the best location filtering would be more of geographical regions like Pacific vs Atlantic slope for example (which is usually stated in a taxon’s publication).

So, as I already mentioned, it only works if the data is there, So it is good in the areas where Kew focuses/focused and for which they have written floras as that is the kind of digitized data POWO is primarily for (and for which Kew has the copyright). Though of course any digitized data could be added.

What you can search for is on the help page: https://powo.science.kew.org/search-help

For some families like grasses the data is very structured: https://powo.science.kew.org/taxon/urn:lsid:ipni.org:names:417047-1/general-information#descriptions Just click on the GrassBase description. At the bottom of the page are the sources of the data displayed.

Thanks for that extra info. It’s nice to know POWO has these capabilities, but I still feel this is a long way from what @letiziaw and I are looking for.

  • Appears to be dependent on whether POWO has incorporated description data from another source
  • Entirely dependent on the wording chosen by the original source. No “controlled vocabulary”, so we’re dependent on the wording used in the source.
  • Geography limited to the fairly coarse regions chosen by POWO.
  • For numerical characters, only supports exact match, not ranges. So, if I’m looking at a grass with awns 35 mm long, and the description of the species says the awns are 32–45 mm long, I won’t get a match unless I search for 32 or 45 mm.

POWO is not the right platform but I would prefer to see people contribute to existing platforms and as you discovered, more is already possible then you think. As I mentioned, this is what WFO wants to do. Apart from the POWO data, it has more New World descriptions and specimen occurrences. https://www.worldfloraonline.org/taxon/wfo-0000279961#distributionMap So rather than reinvent the wheel, try and build on it to get what you need.