I built an LLM-assisted database of plant name etymologies — is this a useful approach?
Hi everyone,
I recently created a web project called Etymon Plantae, a searchable database for exploring the etymology of plant scientific names.
https://lightbox-archive.com/etymon-plantae.html?lang=en
The database currently focuses on plant taxa found in Japan and is intended to cover the Japanese flora comprehensively. You can search by a scientific name, genus name, or specific epithet and see an explanation of the possible origin and meaning of the name.
Although the dataset is centered on Japan, many genera and species are shared with other parts of the world, so I thought it might also be of interest to iNaturalist users outside Japan.
What interests me most about scientific names is how much information can be hidden in them: references to morphology, habitat, geography, mythology, people, historical circumstances, and sometimes rather unexpected stories.
However, there is an important aspect of this project that I would especially like to discuss here.
The etymology descriptions were created with the assistance of an LLM.
I am aware that using an LLM for this kind of information raises legitimate questions about accuracy, hallucination, reliability, sourcing, and whether AI-generated explanations are appropriate at all for a resource dealing with scientific nomenclature.
For that reason, I am not only looking for corrections to individual entries. I would genuinely like to hear critical opinions about the basic idea behind the project.
For example:
-
Do you think using an LLM to investigate and explain the etymology of scientific names is a reasonable approach?
-
Can an LLM-assisted resource like this be useful if its limitations are clearly stated, or does the possibility of fabricated or incorrect etymologies make the approach fundamentally unreliable?
-
What level of citation, sourcing, or human verification would you expect before considering such information useful?
-
Are there certain types of etymological claims that should never be inferred by an LLM without a primary or authoritative source?
-
Would it be better to present several possible interpretations when the origin of a name is uncertain, rather than giving a single explanation?
-
If you think this approach is likely to create more misinformation than useful information, would you recommend that I discontinue the site rather than continue developing it?
-
On the other hand, if you think the concept has value, what changes would make it more responsible or scientifically useful?
-
Would links to original descriptions, botanical literature, dictionaries of Greek and Latin roots, or other references substantially improve the usefulness of the database?
-
If the project were expanded beyond the flora of Japan, are there particular countries, regions, or floras that you think would be worth adding?
-
More broadly, what would make something like this genuinely useful to botanists, naturalists, and iNaturalist users rather than simply an interesting AI experiment?
I am genuinely open to criticism, including criticism of the premise of the project itself.
I do not have a strong commitment to keeping the site online if people with more experience in botanical nomenclature and taxonomy feel that presenting LLM-assisted etymologies in this way is inappropriate or potentially misleading. I would rather hear that criticism directly than continue developing something that experts consider fundamentally unsound.
At the same time, if there is a responsible way to build this kind of resource — for example through better sourcing, uncertainty labels, community corrections, or human verification — I would be very interested in exploring that direction.
If you are curious, please try searching for a scientific name you know well. I would be especially interested in examples where the explanation is clearly wrong, questionable, incomplete, or surprisingly good.
Any criticism, corrections, suggestions, or thoughts about the use of LLMs for this kind of project would be very welcome.
Thanks for reading!