Moderator's introduction

And I'd like to introduce the next speaker. But first, a short preamble. Sometimes we come across various digital, and not only, media files. That's more accurate, knowing what the next talk is about. But they come in completely different quality. And how to work with all that, our dear friend Dmitry Boroshchuk will tell us today. Please welcome him.

To start: what the room uses to clean up photo, video and sound

Hi everyone.

Seriously. Let's start with a chat. How often do you come across a video or audio file where everything is bad? Raise your hands. One. Oh, great. And what do you clean it up with? Colleague, I feel your pain. What do you clean it with? Right from your seat.

— Well, give him a microphone. Yes, let me run around. I could use a second person. Well, usually we apply a model that improves quality. It's, um, minimally, I forgot what it's called. It increases the resolution. And often we use neighboring frames in the video to figure out what's visible in one part and in another. And so, using natural intelligence, you get to, say, the license plate you're looking for. You mentioned the AI, but do you file that in the case too? Like, the AI said: here's the plate number. Well, of course not.

— Good. And what about sound? Over here, another colleague. There was a hand. Dmitry Vladimirovich, I saw another hand there about photos. It was raised high. Actually, this question came out of those ordinary conversations that start in the kitchen and begin like this. Well, we pulled the camera footage, took a look, nothing there, and moved on. Yes, colleague. Originally about photos, I wanted to say... With Amped FIVE and other software solutions, but as for sound... Which ones? Well, various algorithms, for example Richardson-Lucy... Photo Expert, yes, great tool.

— As for sound, I enhanced it exclusively with a neural network. Well, in this case, as for enhancing the sound itself with non-AI methods, there's a compressor... Reverb and so on. Well, that's for vocal material specifically, so to speak. Vocal-instrumental, practically. Well, something like that.

— Okay, so, who else has what problems? And what's your problem, colleague? Well, for sound it's Audition, iZotope by noise pattern, if we have a lot of noise. Naturally, we pull it out, or otherwise we just keep listening. I don't use the AI, because it distorts things heavily, and even the transcription is complete garbage. Yeah, I agree. And what's the biggest problem you run into? Can't see anything, too dark, or too bright, noisy, too little... Poor visibility, right?

Look, today I suggest we talk about... Why did I ask about the software? Because most often, at the moment we need it, the necessary... Licensed software is what we don't have. Or it's just not on hand at the right moment, right? Or we simply don't have access to it. And mostly the detectives wrap up the whole investigation like that... Well, can't see anything.

Free tools only, and no AI

And today I suggest we talk about the software products that are freely available to you, openly available. You can download them freely, without cracking anything. We'll only talk about free tools and only about widely available ones. And, essentially, so that all our actions are, as forensics requires, reproducible afterwards. Why can't we use AI? Yes, the AI is great. And most likely, using AI looks pretty much like it does in that TV show.

It's a magic button. But unfortunately, there is no such magic button. And more often than not we have to be a bit of a photo editor, a video editor, and a bit of an audio specialist, to clean up the moment we need in a photo, in a video, or a recording, say, from a voice recorder. We'll be talking strictly about algorithms. The slides will have step-by-step instructions on what to do for cleanup. And this part will follow the questions that field officers usually come to us with.

Video: pick the codec, check the metadata, rescue a dark recording

First question: we got a video, and a standard player won't read it. That happens a lot. It was pulled off some DVR or some specific piece of recording hardware, and a regular video player won't read it. What can we use here? Of course, we need to find the right codec. I'm going to play Captain Obvious a bit here, but without this foundation we won't understand what comes next. Two excellent, long-established video codec packs. K-Lite Codec Pack for Windows, and a universal player that chews through practically 95% of all media content, called VLC Media Player.

If we still couldn't play it back. Then we need to transcode it. Because the file may be something specific, with a specific extension. And the signatures won't directly tell you what to play it with. There's an excellent solution called Shutter Encoder. It's completely free, and it's very convenient. And its main feature is that it lets you transcode any video into any video. Without losing quality. You surely know that a lot of data can get lost during transcoding. And Shutter Encoder is exactly what will let you, at minimum, transcode without quality loss. At most, see some metadata associated with the file that we can then use in our further work.

MediaInfo is another program. Which, as you've guessed from the name, lets us pull out every possible bit of metadata. To check, for example, whether our file was modified on its way from the source to us. When it was made, what it was made with. And lay all that out visually. The next issue that comes up: the video needs to be made readable. But you can't apply lossy compression. So as not to destroy the fine details. We go back to Shutter Encoder, which we talked about earlier. Next. The defense submits a video from a phone. We need to check exactly which phone it was shot on. Another metadata viewer and extractor is Metadata++.

Also a free application. You've surely heard of it. It lets you pull metadata out of practically any file. And pull out as much of it as possible. And thanks to it, you can extract metadata not only from media files. But basically from everything else too. And it's a pretty decent cataloger.

Next problem: the DVR footage was recorded at night. The image is very dark. You can barely see anything. How do you brighten the scene without quality loss? There can be a lot of steps. The first thing that comes to mind is to fiddle with brightness and contrast. And somehow try to brighten it. In 99% of cases that will get you nowhere. Here you need to understand how the picture is formed. First of all, any photo or video image is a set of pixels. If there's no data in the pixels, then you won't pull anything out. So here we'll be working with colors. With the colors that may be too overexposed or, conversely, too dark.

Another free program. Who here does video editing? The young ones. I've seen a lot of young colleagues. It's called DaVinci Resolve. An excellent video editor. TikTok clips are made with nothing else. But on top of everything else, it's also an excellent tool for working on the picture. Next we'll have an algorithm. I won't dwell on it too much. You can try it out yourselves, after all. So we can squeeze into the time slot.

Next point: license plates. Oh, look at all the phones going up. Colleagues, I'll share a link afterwards. This presentation will be much easier for you than watching these slides with my mug on stage later. So, reading a plate number. How? Actually, there are a lot of additional factors we have to take into account regarding how that plate ended up in the image and what was happening on Earth at that moment. The most common thing you run into is blur. Motion blur because the camera was moving. Someone was shooting with shaky hands or on the run. Or the car was moving. Here we'll need to sync up with the picture. Take the neighboring frames and, essentially, try to reconstruct that plate by removing the blur.

On the screen, as an example, there's a program called VideoCleaner. A fairly old, old software product, a forensic one, a free software product, which is probably about 7 years old now. But it still lets you perform some simple operations. You can also do this with that same DaVinci Resolve. And the Mobile Criminalist folks are supposed to release an e-zine tomorrow, a sort of electronic magazine as a little website, where there'll be my article on how to deal with these kinds of things in more detail.

So, VideoCleaner is a great help. DaVinci Resolve is a great thing. Moving on.

A pixelated face, a licence plate, a reflection in glass

In the surveillance camera footage the perpetrator's face is a pixelated blob. Can it be made recognizable for identification? Can it?

Now, I won't talk to you about artificial intelligence here, because most of the time, when we have a 64 by 64 pixel image, at best the artificial intelligence starts fantasizing. You get all sorts of interesting things, but unfortunately they have no connection to reality whatsoever. So we need to try to pull some additional data out of the neighboring pixels. It doesn't always work, but using, again, a fairly old super-resolution method, you can try to do interpolation and at least bring out specific areas that will help us identify the person. Moving on. There's a glare on a window in the frame.

Feels like the show "Sled," right? Which possibly contains the criminal's reflected face or contains some reflection that will let us identify, well, at least the place where this is happening. Actually, yes, a reflection contains a wealth of information, which, unfortunately, for some reason colleagues don't notice. And here, first and foremost, we act as... photo editors, where our main task is not to improve the picture, but, on the contrary, to degrade it so as to pull out that very reflection from where you'd think it couldn't be. So pay attention to those reflective surfaces which, hypothetically, might contain those files.

Just as free and wonderful: GIMP. You've surely heard of it. A great alternative to Photoshop. A free alternative to Photoshop. And the same DaVinci Resolve, so we can pull together, at the very least, the frames with a convenient, or rather a clearer reflection. Next point. Several witnesses filmed a fight on phones from different spots, and we need to somehow reconstruct the events. There are lots of small software products, free, open-source products, that let us play several video files in sync and tie them to the location. But again, so you don't spend ages hunting for them, use that same DaVinci Resolve and lay everything out on the editing tracks, on the timeline tracks, so that later you can sync up on some event, a flash of light or some loud sound, and switch between those tracks during playback to see what you can piece together from the other camera.

Sound: what can and cannot be recovered; filters in Audacity

Well, let's move on to audio. What problems do we usually have with audio? Nobody has any problems?

— Never had any.

— Noise? Right, very quiet, lots of people talking, everyone at once. And, again, something has to be done about it. And let's start with the fact that first we need to understand what we can pull out and what we can't. We'll remove sound that completely masks the speech. There's no way around that. We can't paint in words that are missing. The AI, unfortunately, doesn't work here. You can play around with processing, but we're unlikely to get a reproducible result. Because every time we process it with that same AI, we will, unfortunately, get completely different versions. But we can even out the volume, improve intelligibility. We can, after all, remove those moments when the mic rubs against a jacket lapel.

The most common artifact that gets in the way of recovery. And we'll do all of this purely based on the physics of sound. Let's first understand what can get in our way when playing back some speech. It could be some constant background. As my colleague said, we take... Just a noise sample, remove it, and listen. Yes, that's an option. When, for example, we have a loud air conditioner running constantly in the background, or some monotonous noise can be heard.

It could be some low-frequency rumble. For example, when we try to record something outdoors. When the urban environment adds the roar of cars, a construction site working nearby. It's wind. Actually, wind and the noise from constant rubbing against some surface are roughly the same. It's hiss, it's rain, it's city noise, it's clicks, and it's clipping. When the sound level exceeds everything and our recording equipment simply cuts off the frequencies. Accordingly, we can't get it back. As I said, we use Audacity, a free audio editor, which lets us remove all this interference, all these noises. Again, each has its own filter here, where we can try playing around with removing it.

So, let's try to answer the question, when some audio comes to us. What kind of noise are we dealing with? Is the noise constant or does it change? What is it mostly like? Clicks, hiss? Where exactly on the spectrum is the main interference concentrated? Does it mask the speech completely? For this we go down the path, essentially, of needing to visualize everything we see. We turn on two types of spectrogram, which will show us exactly where to look for those noises, where to look for the voice we can isolate.

Well, and then. Usually the noise is monotonous, it hardly changes, clearly audible in the pauses, the spectral picture is stable. Here everything is fairly simple. As my colleague said, we take Audacity, we cut, we look for a spot where there's no useful signal, we cut out that noise, feed it to the automatic noise remover, and it removes that monotonous noise from the recording, again, under ideal conditions. It rarely happens that the noise is monotonous and, in principle, all these steps are enough to get it done. Next is rumble. The noise feels like vibration, rumble.

When the main energy, the sound energy, is at the bottom of the spectrogram, when the recording is overloaded, when the mic gets touched. A different story. We go into effects, we go into the equalizer, and we try playing with a high-pass filter to remove that very noise. With wind it's roughly the same. When there's wind, because wind and rumble, in principle, give us the same picture.

Next point, when the noise sounds like hiss. Here it's the upper frequency range, which is on the spectrogram, where we can also try to clean it out with the reverse noise reduction for the low...

through a low-pass filter. City noise. The noise changes constantly, sources appear and disappear. The spectral picture is ambiguous. Here we're already working across the whole spectrum. We take the equalizer and try, first of all, to mark the places where our noise is especially pronounced. Which type of noise is especially pronounced. And we try to work around it. Using a graphic equalizer. And, as the final thing, when we need to bring volume swings to one constant level, that process is called compression, which you can also do thanks to this algorithm. And normalize all the audio you had in the thick of it.

That's actually everything I wanted to tell you today. If we happen not to know each other, I have a little hobby, two Telegram channels.

Q&A

And on that channel, most likely, this presentation will appear in a few minutes. Questions? Before the questions, I'll point out that the article you wrote, Dmitry, Vladimirovich, is already up on the portal with the QR code we showed today. So it's already available, basically. And now, colleagues, yes, questions please. I'm looking at your hands. It's as if everyone... No, there are hands. That's usually when either everything's clear or nothing is.

— Good morning. I have a question, probably a rather silly one. If we have a recording where nothing can be made out, and it doesn't matter if it's audio or video? There's an investigation into some case. That recording was cleaned up and submitted as evidence. Can the defense lawyer object and say the recording was altered, or not? I mean, will it be admitted in court as evidence? Or is that a question for the lawyers? If you used the AI, then the lawyer will say: "Yeah, some weird crap, and it made it all up."

Since all the operations we've just been talking about are reproducible, then after the expert testifies... After the expert or specialist testifies, the court usually accepts it.

— Dmitry, thank you for the talk. You're welcome. As a former forensic audio examiner, I can just say that what you've presented is roughly the level of maybe 2006 or 2007. I wish you to pick up more methods. In particular, if our noise has structured harmonic components, they're very easily removed without damaging the speech. There's a little program, sadly the author is dead, called Justiphone. Yes, a wonderful thing. It removes harmonic components beautifully. There's a demo example of a conversation against an organ background. When we have a clearly defined interference, we can subtract it without damaging what lies underneath.

Second, on video. Here too I'd advise colleagues not to trust too much what's visible at first glance, because you didn't cover all methods. Absolutely. And on the practical side, we had to reconstruct a shootout. You probably know. There was a shootout at Wildberries back in the day. And there were about 45 video recordings from different sources: fixed cameras, dashcams, Dozor body cams, mobile devices. And we had to manually stitch the video streams into a wider frame, so that afterwards, instead of switching, we could watch them at once; of course it was hard on the eyes, but we analyzed who was shooting from which side and whose pocket that pistol had been in.

Here's what I want to add. Just an addition to your example. It's simply very good when we don't just switch between video streams, but view them simultaneously, in a sort of multiplexed mode. Thank you. Yes, thank you very much, first of all, colleague. Yes, 2006 to 2008, absolutely, the same VideoCleaner. The question is different. What can we use when we have nothing at hand? And what can we use when we don't really have any special expertise? But it has to be done. One moment. About event reconstruction. One moment. Dmitry Vladimirovich, we're still being filmed. May I? Oh, sorry. On stage, please. Yes. About event reconstruction.

It really is a problem when we have several sources, we have close to a hundred witnesses we've collected video data from, which then has to be placed on a map. And also with elevation, and the viewing angle, and tied to the timeline. A little teaser. My team is right now preparing exactly such a tool for reconstructing events from video data with geolocation to the terrain. It will most likely be free, and you'll be able to use it. Something like that.

Dmitry Vladimirovich, thank you very much. Unfortunately, we're out of time. Let's see Dmitry off with applause.

And now I probably have some not-so-good news for those who are with us online, because for today our event is over. But only for those who are online. The next parts won't be streamed, and we'll be back with you tomorrow at 11 in the morning. But for those here in the hall today, we're moving on to our, let's say, more closed-door, more hands-on part of the event.

[End of the day 1 stream. What follows was not streamed: the talk by Ekaterina Mingaraeva (SEUSLAB, 13:40–14:15 in the program), the closed-door part of day 1 (after 14:20) and the opening of the 2021 time capsule. The next line is already the opening of day 2, 4 September.]