Sounds - I want to screen out mobile phone captures

Really, I get it. the app are convenient. But seriously, if you CBA to trim humans, filter audio a bit, normalise so I can actually hear something, then I’d like to be able to at least screen out your random capture so I can hear something. So if I could screen out mobile phone captures on the ID then I could focus my effort of IDs that I can a) hear and b) recognise.

I’m not saying don’t have mobile phone audio. Just let me block it out if poss. I’m just not prepared to listen to any more of this garbage.

Sounds to me like this would be a feature request if no one pipes in with an answer to your polite request for help. Welcome to the forum.

It has less to do with the technology and more about a user’s skill and understanding of capturing good, identifiable media. This would be the same as wanting to filter out observations with cell phone camera photos and only viewing observations with photos taken with digital cameras.

Quality recordings can be taken with cell phones: https://www.inaturalist.org/observations/278853592

Thanks, yes, probably a feature request. I’m not going to hit people with un-normalised audio, I will trim extraneous garbage out and I will EQ traffic noise out . I can’t eliminate aircraft, but I will try and make the wanted signal more than the rubbish. If you want me to try and ID your stuff, then at least make it audible :wink:

Also, it’s never really a good idea to call people’s observations “garbage”.

as in I trim handling noise :wink:

This “garbage”.

I have tried to ID a lot of audio. Signal quality does matter. I absolutely accept some people with better hearing may be able to detect signals I can’t cope with. All I want to be able to do is screen out recordings people make with mobiles, because I CBA to listen to stuff that I would torch for signal quality. It should be easy enough, and you’d get more ID from me. It’s saving my effort, I absolutely am not saying mobile recordings have no value. I’d just prefer not to allocate effort to them.

Fair cop :wink: Too much stuff I couldn’t hear, but yes, my bad

Most people can’t afford fancy recording equipment. I think recording with your cell phone is fine as long as you tinker with it in Audacity (free) until it sounds good. (edit: and are relatively close to your subject)

I’m hoping some day we will get a built in (iNat generated, so standardized) spectrogram so I can see how loud/quiet or noisy a file is before I listen to it.

Can you direct people to a good resource on how to do that? I don’t record audio with a cell phone, but I don’t use fancy recording equipment either, so it’s probably of comparable quality.
I have Audacity on my laptop, so I could use it if I knew how. I’ve been afraid to try because I really have no idea where to start. With image editing, I can see the effects and determine which ones are an improvement. I don’t know how to do the same with audio. It seems like it would be easy to accidentally distort the sound in a way that makes it unidentifiable, and I don’t know how to recognize where that line is.

Plant, animal and fungi captures suffer from the same issue, just mark them as reviewed and don’t interact with it if you choose to

Most people uploading difficult-to-hear audio files are probably not doing so intentionally. Quite likely they have no idea how to edit audio files – it’s not something the average person has much experience with, and phones don’t necessarily have built-in intuitive editing tools for audio the way they do for images (often they don’t even have a dedicated standard app for just recording). In addition, new users often have never consciously thought about how something is identified and how they can provide media that makes it easier to do so.

I think what is most needed would be tutorials on making and editing recordings to be provided somewhere prominent and easily findable on iNat as part of user onboarding (though we have been asking for better onboarding for so long that I am doubtful that we will see any meaningful progress on this front anytime soon).

Many people would also probably find it useful to have an embedded tool for basic audio editing, but given the relatively small proportion of observations on iNat that are based on audio, it would likely be necessary to weigh whether the benefits are worth the costs (development/hosting etc.)

The easy wins are trim out extraneous stuff, usually beginning and end. Often high-pass filter at ~ 300Hz if the subject is a bird, obv not if it is an elephant :wink: After that normalise level, typically to about -3dBFS. This guy shows how https://www.youtube.com/watch?v=Bwe6tawueOY

The problem with qualifying audio ID is that it’s a horrible experience on iNat. Qualifying audio is always going to take longer than images, because you can scan through a lot of images at a glance. on here you have to open each sound . if you wick up a quiet sound, the next one with handling noise and stupendous wind rumble is actually painful to hear. Audio normalisation can be done programatically. Some places preview as you move the mouse over the file, which would be a help. A spectrogram like https://xeno-canto.org/ or a waveform a bit like https://freesound.org/ would at least visually show the stuff that’s all handling noise and rumble. https://aporee.org/maps/ does neat things with maps

I take the point, though it was hard to remember after the twentieth load of wind rumble and handling noise, interspersed with stuff so highly compressed the top end of a robin’s song was lost and the ID was more to do with cadence than pitch. Sure, someone one day may capture the sound of a relict population of Ivory-billed woodpeckers with a phone. I’m not saying get rid of phone captures.

All I wanted was to screen them out. I could screen by county and by birds and by has sounds, and thought I could give something back by helping with the more obvious sounds. But I really can’t face any more untrimmed mobile phone recordings. With images its easy to pass a shot where there’s a thumb over the lens, with audio screening is much harder with the current interface.

I would like to add here that there are some benefits of “garbage” audio. At least I hope that we get at some point an auto-ID for audio too, and the more realistic the training material is, the better the model will be able to handle real-world audio recordings.

Just like a bad picture, sometimes nothing can be done, but personally, I get great pleasure out of ID’ing something difficult, even if it is difficult because the quality is bad

With time identifying, you’ll see who is submitting higher or lower quality audio. Then you can choose to identify observations from selected observers, and go into others and mark them all as reviewed so they don’t show up for you again. Or edit your Identify filter to exclude certain users.

I have been trying to help with bird audio identifications this past summer. I discovered that you can right-click on a recording in iNat and download it. Then, using Audacity (a free application) you can enhance recordings (apple files, WAV files and MP3 files). You can somewhat filter out noise and also increase the volume fairly easily. There are a lot of bad (phone) recordings out there especially from people who use Merlin. My biggest issue is with the ones that are 1 to 3 seconds. There almost always isn’t enough information recorded. But, there are also a lot of very good recordings on phones that just need a little bit of cleaning up. So, it’s worth the effort to download them and work with them a little bit. I also sort by Ascending to start with the older recordings.

And, thanks to a wonderful iNat user, wweellll, you don’t even need to download things and use Audacity. You can use this persons free website that was created to help clean-up recordings:

This works for videos and audio files.
I tried it with a very quiet Baltimore oriole recording I made.
The recording was noticeably improved. It was brighter. The surrounding bird songs were clearer. I think I “discovered” two or three birds in it that were too quiet to hear in the initial recording.

https://normalizer.pages.dev/

https://forum.inaturalist.org/t/quick-and-easy-tool-for-cleaning-up-audio-observations/65427

I don’t have any resources to direct you to. I just found tools that I thought were easy to understand/use.

When you have it set to view as a spectrogram you can see things you don’t want and the results of changes. I also try to do long recordings so I can find the best part to upload. I’m trying to avoid things like dogs barking, roosters crowing, car engines revving, and very loud Carolina wrens and Mockingbirds. Also sometimes people talking, but usually I’m by myself or with one other person whom I’ve asked to stop while I’m recording.

I just mess with the amplify slider (and listen to preview) until it is easily audible from my laptop speakers with the volume at 100%. If I can’t hear it unless I have earbuds it, then it’s too quiet IMO. If you try to boost the volume too much though it can start to sound a bit distorted. I don’t know how to tell you what that line is though. You can upload multiple files with different levels of amplification (or noise filtration) and indicate this in the notes section. Or even include a copy of the original file. If I do this, I will put the edited version as the first file and the original as the second file.

Other people use high/low pass filters, but I find those difficult. I prefer using spectral delete because I use my cursor to kind of crop out what I don’t want. Usually noise from insects is higher pitched than birds and frogs that I want to record and I can use spectral delete to remove them (or to remove the birds/frogs if I want the insects to be the focus of the recording). If they overlap I have to use the noise filter which is more difficult. Low pitches that I’m deleting are usually wind, cars, or airplanes. Unless I’m recording something like doves or bullfrogs, usually there isn’t pitch overlap.

Example spectrograms of some crickets

Original:

After amplified by 25 dB:

After spectral delete of low pitch wind sounds:

I think I amplified it a little more than 25 dB back when I originally edited it, but I was just doing this quickly to get some screenshots.

Thanks for this. As mentioned in another comment, I try to do bird audio identifications. The high-pass filter does get rid of things like wind and other low sounds, and really helps with bird songs which are high-pitched. The normalize also helps clean up a recording.