# Media forensics: improving photo, video and sound Dmitry Boroshchuk · Beholderishere.Consulting MOSCOW FORENSICS DAY ’26 · Day 1 — Thursday, 3 September 2026 · Scheduled 13:05–13:35 · In the recording 01:52:11–02:20:07 Talk summary · https://2026.moscow-forensics-day.workers.dev/en/summary/05-boroshchuk Transcript: https://2026.moscow-forensics-day.workers.dev/en/transcript/05-boroshchuk · Slides: https://2026.moscow-forensics-day.workers.dev/en/slides/05-media-forensics · Watch from 01:52:11: https://youtu.be/WuMIv5sFPRs?t=6731 --- ## In brief A practical set of free tools for the cases where a recording "won't read", is too dark, blurred or noisy — and a principled refusal to use AI: not because it works badly, but because the result is not reproducible, and reproducibility is what a court needs. Video: picking the codec, lossless transcoding, metadata, brightening through work with color, removing blur using the neighboring frames, super-resolution for a pixelated face, reflections in glass, syncing several recordings on the editing tracks. Sound: first an honest list of what cannot be pulled out, then a breakdown of the types of interference by spectrogram and the filters in Audacity. The talk opens with a poll of the room and is tested by that same room: a former forensic audio examiner says outright that the level of the methods is "2006 or 2007", and delivers what was not in the talk. ## Key points - The opening — a poll of the room: what do you clean up bad video and sound with. The answers: an upscaler model and piecing a license plate together from the neighboring frames, Amped FIVE and deconvolution algorithms for photos, Audition and iZotope by noise pattern for sound. Asked whether the output of an AI is filed in the case, the answer from the room: "Well, of course not". - A remark from the room against AI in sound: "it distorts things heavily, and even the transcription is complete garbage". - The problem stated: the licensed software needed is usually not on hand or there is no access to it, so "mostly the detectives wrap up the whole investigation like that… Well, can't see anything". - The frame of the talk: free and widely available tools only, "you can download them freely, without cracking anything", and only what gives a reproducible result — otherwise it is not allowed in forensics. AI is excluded: "there is no such magic button", you have to be your own photo editor, video editor and audio specialist. - The video will not open in a standard player (shot by a DVR or some specific device): the codec packs K-Lite Codec Pack and the VLC player, which "chews through practically 95% of all media content". - If it still will not play back — transcode it in Shutter Encoder: free, it "lets you transcode any video into any video" without quality loss, and it shows the file's metadata. - MediaInfo — the full metadata: whether the file was modified on its way from the source, when and with what it was made. The same task, "the video needs to be made readable. But you can't apply lossy compression. So as not to destroy the fine details" — again Shutter Encoder. - To check which phone a video submitted by the defense was shot on — Metadata++: it pulls metadata out of practically any file and works as a cataloger. - A dark night recording: fiddling with brightness and contrast "in 99% of cases… will get you nowhere". You have to understand the nature of the frame: if there is no data in the pixels, there is nothing to pull out; you should work with the colors — the overexposed and the darkened ones. The tool — the free DaVinci Resolve. - A license plate: the main obstacle is blur and "motion blur" (camera shake, running, the car moving). The method — sync up with the picture, take the neighboring frames and reconstruct the plate by removing the blur. - The tools for this: VideoCleaner — an old, "probably about 7 years old", free forensic product, still good for simple operations, and that same DaVinci Resolve. A detailed walkthrough is promised in an article on the anniversary MFD portal. - A face as a pixelated blob: with a 64×64 frame the AI "starts fantasizing… unfortunately they have no connection to reality whatsoever". Instead — super-resolution and interpolation over the neighboring pixels: "it doesn't always work", but it can bring out areas useful for identification. - Glare on glass: a reflection carries a lot of information that "for some reason colleagues don't notice", and the task here is the reverse — "not to improve the picture, but, on the contrary, to degrade it so as to pull out that very reflection". The tools — GIMP as a free alternative to Photoshop and DaVinci Resolve. - Several witnesses filmed a fight from different spots: lay all the recordings onto the editing tracks of DaVinci Resolve, sync them on a common event (a flash of light, a loud sound) and switch between the tracks. - Sound begins with the limits of the possible: sound that is completely masked cannot be recovered, "we can't paint in words that are missing", the AI is no good — "every time… we will get completely different versions". You can even out the volume, improve intelligibility and remove the mic rubbing against a jacket lapel — "the most common artifact". - Next — "purely based on the physics of sound". The types of interference: a constant background (an air conditioner), the low-frequency rumble of the city and of a construction site, wind (which by its picture matches rubbing against a surface), hiss, rain, city noise, clicks and clipping, when the equipment cuts off an overloaded signal. - There is a single tool — the free Audacity, "each has its own filter here". - Diagnosis before processing: is the noise constant or changing, clicks or hiss, where on the spectrum the interference is concentrated, does it mask the signal completely. For this, two types of spectrogram are turned on. - Monotonous noise: find a spot with no useful signal, cut out a noise sample and feed it to the automatic noise remover — but this works "under ideal conditions. It rarely happens". - Rumble, vibration and touches on the mic (the energy at the bottom of the spectrogram): the equalizer and a high-pass filter; with wind it is the same. Hiss — the upper range, removed with a filter. - City noise is the hardest case: the picture changes constantly, sources appear and disappear, so you work across the whole spectrum with a graphic equalizer, marking the most pronounced interference. The final steps — compression (evening out the volume swings) and normalization. - The presentation is posted to the speaker's Telegram channel right after the talk; the article is already available on the anniversary MFD portal. ## Tools, artifacts, technologies - **Video**: **K-Lite Codec Pack**, **VLC Media Player**, **Shutter Encoder** (lossless transcoding), **MediaInfo** and **Metadata++** (metadata), **VideoCleaner**, **DaVinci Resolve** (color work, editing tracks, syncing several recordings). - **Photo**: **GIMP**; the methods — super-resolution and interpolation, removing blur using the neighboring frames, working with reflections by deliberately "degrading" the frame. - **Sound**: **Audacity** — two types of spectrogram, cutting out a noise sample and the noise remover, an equalizer with high-pass and low-pass filters, a graphic equalizer, compression, normalization. - **Named in the remarks from the room**: **Amped FIVE**, deconvolution algorithms (Richardson-Lucy), Photo Expert, **Adobe Audition**, **iZotope**, **Justiphone** (removing harmonic components), Dozor (a source of video recordings). - **Types of interference**: a constant background, low-frequency rumble, wind and rubbing, hiss, rain, city noise, clicks, clipping. ## Legal and organizational context The main legal motif of the talk is reproducibility as a condition of admissibility: everything shown was chosen precisely because the steps can be repeated and demonstrated. The direct question about the defense lawyer's position gets this answer: if AI was used in the processing, the defense will say that "it made it all up", and since the actions described are reproducible, "after the expert testifies… or specialist testifies, the court usually accepts it". The remark from the room that the output of an AI is not filed in the case matters separately. The organizational background is the shortage of licenses in the field, because of which an examination often does not begin at all. ## Questions from the audience - **1** (name not given): if a cleaned-up recording was submitted as evidence, can the defense lawyer claim that the recording was altered — will the court admit it? → If AI was used, the objection will be well-founded; with reproducible actions the recording is usually admitted after the explanations of an expert or a specialist. - **2** (a former forensic audio examiner, name not given): direct criticism — "what you've presented is roughly the level of maybe 2006 or 2007", and advice to pick up more methods. The specifics: noise with a harmonic structure is removed easily, without damaging the speech — the program Justiphone can do that, and it has a demo of a conversation against an organ background: when the interference is mathematically defined, it can be subtracted without touching the signal underneath. On video — "not to trust too much what's visible at first glance", not all the methods were named. From practice: in reconstructing the shootout at Wildberries they collected about 45 recordings from different sources (fixed cameras, dashcams, Dozor, mobile devices) and manually stitched the streams into one wide frame, so as to watch them at once rather than switch between them — that way you can see who is shooting and whose pocket the pistol had been in earlier. → The answer: agreed, "2006 to 2008, absolutely", but the talk is about something else — "what can we use when we have nothing at hand" and when there is no special expertise but the job has to be done. - **3** (a continuation of the same subject): reconstructing events from many sources — "close to a hundred witnesses", tying them to a map, elevation, viewing angle and timing. → A teaser: the speaker's team is preparing exactly such a tool for reconstruction from video data with geolocation to the terrain, "it will most likely be free". ## The speaker's position The position is pragmatic and deliberately down-to-earth: the talk is not about the best methods but about the minimum available to any operative without licenses and without training. The refusal to use AI is justified not by quality but by the procedural unfitness of a result that cannot be reproduced, and that is the strongest idea of the talk. The speaker accepts the criticism from the room without arguing and immediately restates his task. The limits are spelled out regularly ("it doesn't always work", "under ideal conditions", "if there's no data in the pixels, then you won't pull anything out"), there are no promises, and no product is being sold. ## Quotes - "…mostly the detectives wrap up the whole investigation like that… Well, can't see anything". - "It's a magic button. But unfortunately, there is no such magic button". - "…when we have a 64 by 64 pixel image, at best the artificial intelligence starts fantasizing… unfortunately they have no connection to reality whatsoever". - "…our main task is not to improve the picture, but, on the contrary, to degrade it so as to pull out that very reflection". - "…what you've presented is roughly the level of maybe 2006 or 2007".