Day 1 — Thursday, 3 September 2026
Day 1 opening
Dmitry Yankovoy — moderator (MKO Systems).
Housekeeping and the venue
Good day. In just a couple of minutes we'll be starting our conference, so I invite everyone who's currently in the coffee-break area to join us in the hall. And while our friends are making their way in, let me go over a few technical points about the venue itself. So, this here is the presentation area. As you saw in the program, there's going to be quite a lot today. A bit further on is the coffee-break area. During every break there'll be tasty tea, coffee and other snacks for you there. Besides that, if you go just a little bit further, that's where our booth area is, where the conference's partner companies have brought a lot of their developments, because some we haven't seen in two or three years.
Besides that, let me tell you about the entrance and exit. You all came in today through the main entrance, by elevator. No more than six per elevator. The venue staff have probably told you that already. Still, I realize there are quite a few smokers among us, so we don't overload the elevator during breaks, we've opened another entrance. That entrance is on the left side — over on the left there's the restroom, and that exit a bit further on. Also, that's where we go out, but the entrance still stays centralized, just the one, exactly where you came in. That's basically it from me on the technical points, and I think we can get started with our conference.
First, let me introduce myself. I know many of you have seen me before, many have known me a long time, but my name is still Dmitry Yankovoy. Today, following the good old tradition of MKO Systems, I'll be moderating our event today. I'm glad to welcome those who are here in the hall today, and glad to welcome those joining us on the online stream. And let's move on first, besides me, to introducing the company. The main organizer of this event is MKO Systems, a leading Russian developer of software for computer forensic examination of mobile devices, cloud services, computers and much more.
We also run our own training in digital forensics. Today, by the way, we've prepared quite a lot of surprises for you, because our conference has been held since 2016. Imagine, back then it had a slightly different name, but still, the first seeds were planted right then. I think there are people in the hall right now who were at that very first conference, and yet ten years have passed. And honestly, it's really nice to see so many people in the hall who over this time have become our friends, because our conference is really not so much about the talks, about learning something new, as it has become one big platform for people from literally all over the country and different fields.
Because originally the conference was intended only for law enforcement staff, then we changed the format a bit, and every year we experiment with that format. And by the way, this year we'll have experiments too, which, so to speak, will break with our traditions.
The tenth MFD's format: a 2021 time capsule instead of the roast
As you know, in the final block of our event we always hold a roast, but today there won't be a roast. Back in 2021, at the conference held around these same dates near Komsomolskaya metro station, we sealed a time capsule, back in 2021. We haven't opened it since, but today that's what we'll do, and discuss how the experts of that time, and other researchers too, saw digital forensics five years on, and we'll go over all those notes with our experts. So overall the format stays the same, that is, a space for live discussion, but there won't be a roast. So what do we do about bugs?
I know they'll show up either way, I mean, there's no getting away from that, Only those who do nothing never have any. Look, throughout the whole day, as I already said, our team will be working in the booth area. You can come up to the guys, tell them everything, show them everything. They might even help with something right away and advise you, if that's possible. But again, as our tech support always writes: send us the logs. And that's true, honestly, because poking around blind is a tricky business. Still, besides being able to leave feedback there, you'll also be able to pick up some small souvenir prizes there.
So, I've run through the technical points, covered everything. I won't say anything about phones, you all know that already.
Valeria Vakhrushina (MKO Systems) — “MK Bruteforce: a special edition for the tenth MFD”
Scheduled 10:15–10:40.
Moderator's introduction
So let's open today's conference. And for that I invite our marketing director, the wonderful Valeria Vakhrushina. Let's welcome her with a round of applause.
Conference opening: partners and ten years of MFD
— Good afternoon, colleagues. I'm very glad to see you all in the hall today. There are a lot of us today, and today will be interesting.
And my clicker, as usual, isn't working.
Right, oh well, technical hiccups. I'd also like to introduce the partners of today's conference. Thanks, Zhenya. Dmitry, who's taking part with us today at the 10th MFD? Who's supporting us? Today we're supported by companies such as Account Best, ELETEK, Ester Solutions, Amplicom, ACE Lab and SearchInform. But I deliberately didn't do a teaser, because I know the guys will tell you everything far better than me. The guys will tell you everything, but again, besides our booth, where you can talk to our specialists, there are our partners' booths too, where you can ask them tricky questions and find out what's new. And of course, look for them after the talks.
It'll be interesting, I hope. I have a question for the hall. Was anyone here at the first MFD? We're ten today. I'm curious, was anyone there?
Wonderful. So they exist, the ones with us all ten years. Great! Well then, let's start our talks.
MK Bruteforce: new hashes, faster scrypt, dictionaries, plans
Excellent. And we'll start with our MK Bruteforce. I think you're all familiar with it. I hope you've known it a long time. Zhenya, my clicker isn't working.
One more slide.
Wonderful. So, here we go. I think you're all familiar with it. It's our product built on hashcat. I think those of you at the spring conference remember we already updated to hashcat version 7.0.
Overall, what Bruteforce can do you've all known for a long time, I think. But let's look at what new things we've added. These are basically all the hashes we currently support. I urge you: if something's missing, some applications, some hashes that hashcat supports, or that aren't supported at all, you can come up to me any time, or to the guys at the booth, and request it. We'll definitely try to add them. A bit later I'll tell you what's in development now. So, what's fresh? We've added Samsung Smart Switch Backup for you. Telegram for macOS, the passcode for the desktop version, and Threema.
Every time, users ask me where to get the hash. Just in case, I'll show you: we load it straight into the Bruteforce UI. With Samsung Smart Switch, we load two files. Make sure they're from the same backup. That's important, otherwise something may glitch and it'll recover some wrong password. It's exactly the same with Telegram for macOS. You can paste the hash or upload the file so it gets recognized.
And Threema works exactly the same way. Let's see what else we've updated. We've improved the scrypt algorithms, specifically for Android physical images, both FBE and FDE, Apple Notes, Huawei HiSuite backups, Threema and Telegram for macOS. That means everything should now be recovered somewhat faster. Our measurements showed that in some places the speed increased tenfold, in others it's not so rosy, of course. FDE won't crack fast, as all of us who use it know.
But still, try it, test it. I hope everything works for everyone. Again, if anyone has problems, write to our support or write to me directly, catch me in the hall. What else have we done? We all remember that MK Bruteforce runs both on GPU and on CPU. But a lot of people get confused. I often get screenshots saying, look, I have one graphics card and it shows up twice. Note that this depends on which drivers you have. So with this screenshot from my laptop, I have both the CUDA driver and the OpenCL driver. That's why the graphics card shows up twice. In the next release, I promise, the graphics card will show up once, and there'll just be a driver switch.
But in any case, I recommend everyone use OpenCL. It's usually a bit faster, but you can test that yourself on your own device and see which is faster for you. Also, at your request, dear users, we added showing every graphics card separately, so you can see the temperature breakdown. You asked us to add that to the UI. And we added a notification that the attack has been paused on reaching a critical temperature, so it's clear what happened to me, what kind of bug this is, why it's paused, why it isn't working. It all works, everything's fine, we're just trying not to burn out your devices.
Okay, what else did we add? We added a dictionary library. Now you can get in through an interface like this. There are preinstalled dictionaries. You can see they can't be deleted, you can't do anything with them. Here you can add new dictionaries, delete them, edit all of this for yourself somehow. But most importantly, there were a lot of requests to start supporting large dictionaries. So now you can safely load a 20-gig dictionary, a 40- or 100-gig dictionary. It will work, everything's fine, now it's actually possible. The only thing is, keep them where you won't accidentally delete them. Because they're saved here in the interface, and if you edit something, it will sync seamlessly.
So if you deleted a dictionary or wiped something out of it, there could be a problem. Keep a close eye on that. Plans for the future. I promised teasers. What are we preparing now? We're preparing the KeePass application. It'll be out soon. We're preparing the Kims application. It'll also be out soon. And there were a lot of requests for password recovery for Windows Hello. That's in research now, but I think within the next six months, that is, a release from now, give or take, we'll be able to add it. The team is working really hard to make sure everything works for you. So if you have any wishes, if you have any suggestions for improving Bruteforce, we're always open, we'd be glad to hear it.
For those who haven't used it for those who for some strange reason don't have Mobile Criminalist, first, you can come to our booth and ask the guys what it's for. Maybe they'll even share a demo. And second, you can follow the QR code. There's a free version of MK Bruteforce. I know, I haven't updated it in a while. The update will come with the next release of Mobile Criminalist.
All the fresh, tasty new stuff will make its way in there. Of course, it's a slightly cut-down version. It supports fewer hashes than the Bruteforce built into MK. But still, you can use it. It's there. And if it has any bugs, we try not to code them in. But in any case, write to our support too. We'll definitely keep track and improve it. Well, I'm quick today. If you have any questions, I'm ready to answer them.
Where's my Dmitry Yankovoy? With the microphone.
We lost our moderator. He's been found. He wasn't lost. Guests keep arriving, pulling me every way. So, colleagues, I'm looking for raised hands, right? I think they'll grill me later, I know from experience, they'll grill me in the hall.
The 10th-anniversary MFD portal: articles, video archive, a training exercise
— All good overall, but in that case I suggest the next slide. I'll give you the great honor of telling everyone about the first gift we've prepared for today's conference. Then I'll add to it.
— Well then, I suggest you all scan this QR code, there won't be anything indecent there, I promise. It's a small information portal we opened in honor of the 10th Moscow Forensics Day, which we're holding with you today. You need a short registration, and I promise no spam mailings will come to you after that. What have we prepared for you there? Our friends, both forensic and infosec specialists, wrote several articles there. I hope you'll find them interesting to read. Check them out. If, again, you have any wishes, maybe something is missing, maybe someone wants to take part and write something too, some article on their own or together with us, you're welcome at our booth.
We'll be glad to work with you. But that's far from all. Because it's on this very information portal that we finally got around to structuring all the information we have. You won't find our videos on Rutube anymore, but if you're a client of MKO Systems, all you need is a short authorization. As soon as you register on the portal, it'll all be spelled out, what you need to do. I won't dwell on that now. And then you'll again have access to fresh materials, to a large number of manuals, to videos and everything else. But in honor of our tenth event, my colleagues and I thought for a very long time about what interesting and cool thing we could do.
And so we implemented the ability, this will actually be of most interest, probably, to universities, to complete a challenge. So now on the portal you'll also be able to find a challenge. There's a short preamble to the challenge there. A test image is already ready, easy enough to download. And you'll be able to actually solve it. Then it'll say what we want from you. You don't need to upload anything. But you can solve this case, write us your answer, and our colleagues check it. The only thing I'll say right away: I know that before our conference, already underway today, 50 people had registered on the platform.
You won't be able to register with the same email, unfortunately, so down below there'll be a small button that just says "Log in". So you log in with the credentials you already entered on the site earlier, and that's it, it all shows up. If you scan this QR code here, the articles will show up on top of that too. So go ahead, give it a try, write to us, and if anything comes up, we're always here in touch.
Vyacheslav Chikin (ACE NPP, ACE Lab) — “Mobile devices with a dedicated cryptographic processor”
Scheduled 10:45–11:15.
Moderator's introduction
— And now I'll invite Dmitry to announce our next speaker. Yes, thank you very much. Can I get the clicker?
Evgenia Valeryevna, the clicker isn't working again. Could you switch the slide for me, please? Otherwise the magic won't happen. There. So, right now we'll have a presentation from the company ACE Lab. But let's talk in general about what it'll be about. Because after all, today we have a conference specifically on digital forensics. And modern mobile devices increasingly use various hardware keys, access code checks and other critical operations. That is, protection is now built not only at the level of important mechanisms, but it's now moved into a separate, isolated cryptographic processor.
And here a completely different set of tasks arises for specialists. How such protection is designed, what limitations it imposes, and what methods let us work with such devices in practice today. Please welcome the speaker from ACE Lab, Vyacheslav Chikin. Let's give him a round of applause.
Vyacheslav, the clicker.
Talk
Good afternoon, I'm glad to see you all today. And we'll talk about why it isn't possible to work at full capacity with such modern processors as Kirin, Exynos, maybe modern Qualcomm. It's because they have a device, a full-fledged processor that handles cryptography. And we'll try to conduct a little research.
So, this is what a crypto chip looks like on the processor board. We'll look at the cases where the crypto chip isn't built into the processor but is placed separately on the board. Clearly, for our research that's much more convenient. Here we can solder in, see what exchange is happening, and draw some conclusions. Today we'll be looking at devices based on the MTK and Unisoc processors.
Namely, on Unisoc that'll be the Samsung Galaxy Tab A8, and the crypto chip is connected to the processor via the SPI interface, and the Samsung Galaxy A06 on the MTK processor, and here the crypto chip is connected via I2C. Accordingly, here we have two different generations of crypto chips.
So, what does a cryptographic processor actually represent? In fact, it's a full-fledged device with everything that entails. It has its own core, which runs at a fairly decent frequency, which allows it to do a ton of cryptographic computations. And not only that. There are built-in cores, accordingly. It even has its own GPIO. It's a full-fledged device that we're dealing with.
This is for MTK. The only thing here is that it's I2C. But on the other hand, there's so much memory here you could fit whole operating systems in it.
What the key is made of and why you cannot peek at it
So, let's get started. Namely, what interests us. This slide shows the scheme for generating the key for the file system space. Accordingly, in order to get a key, some hash, we need the user's password. Also some kind of hardware uniqueness contained in the processor. It could be a key, a hash. Each manufacturer defines this differently. Certain constants, again, can vary.
A 4-kilobyte blob, secdis, that's also arbitrary. The state of the phone, locked, not locked, that's also taken into account. So if you're thinking somewhere that now we'll unlock the phone and somehow gain access to it, that doesn't work. Because in that case the encryption keys change at once and all the data effectively disappears. Also the so-called color. This is, essentially, the state of the phone.
That is, whether it's unlocked, not locked, what firmware is installed. That is, if the firmware isn't verified, the state changes from green to yellow, and again, we lose the data. And also some so-called ROT. This is data embedded by the manufacturer. Most often it's a public key and a hash of it. And here one more ingredient appears. I called it knox_nonce, some block of data contained in the cryptographic processor.
Okay, well, if some block of data is in the processor and an exchange takes place, we'll try to sneak a look at it over SPI and see what we get. That is, as always, we take the phone, look for the pads, solder, connect the analyzer, and what will we see.
Well, to start, let me tell you what exactly we're going to be up against. If it were all so simple, cryptographic processors wouldn't exist. If we just said: give us the key, and it said: here you go, guys. But that's just no fun. So, to start, the first problem is we need to turn the processor on. Samsung made it clever, it only turns on for a transaction. That is, we want to ask it something, the main processor turned it on, talked with it and says: okay, don't get in the way of my work.
Then, after we've turned it on, the processors have to shake hands with each other. To say: I'm Vasya, you're Petya, you're Petya, I know you, you're saying everything right, let's keep working. After this, the crypto co-processor sends a random number to the main processor. The main processor encrypts it with the established key. And this thing is then verified by the crypto chip. If all is well, the next stage happens.
A secure channel opens, session keys are brought up. And, essentially, the rest of the data we see as just garbage for us and some information underneath it, meant for the processor. At first glance this scheme looks fairly protected, but let's try to look for weak spots in it. So, essentially, this is the native option. We solder in, watch the exchange, try to find something. But even if we sneaked a look at the challenge, after the secure channel came up, we see garbage.
Accordingly, we need not just to see the packet exchange, but also to somehow wedge into controlling it.
Accordingly, if we can't observe it, let's try to desolder the processor somewhere, control the power, and try to go down this path. But here this option doesn't suit us, because we don't know the encryption key. Let's call it that, though there are session keys too.
Accordingly, we have to learn to control the main processor, to somehow get to the data. Okay, right now we're looking at it in theory. Suppose we intercepted control of the processor. Good, we control the processor. Somehow, let's say, we found the key in the processor and repeated the exchange over SPI. Accordingly, we have the key. But the next problem is supplying power to the crypto-processor. You can feed it head-on, bypassing everything, by soldering on some wires. But we might fry something.
So we'll have to deal with this problem inside the processor too. Well, let's say we solved that one too. Then here we can already, by taking apart the transaction, knowing the session keys, look at Knox. knox_nonce. And by adding it, this missing link, to the previous scheme, offload the brute force to an external computer. So here we have GPU power, CPU power, and clusters, so here we're already working at full scale.
In principle, this process turns out to be pretty high-level, but it has its vulnerabilities and the attack can be carried out.
Galaxy A06 on MTK: the crypto chip as a gatekeeper, brute force only on the phone
Now let's move on to the next phone, the Samsung Galaxy A06. Here the difference isn't just that it's controlled over the I2C bus. You can look at the I2C bus, and I'll say even more. We'll see the exchange. It isn't encrypted here. Let's say the trick here is something else entirely. Roughly speaking, it's in forming what's called the key space.
If in the first case we only needed, let's call it, a secret blob, and we calmly offload the brute force outside. Here the crypto-processor is directly involved in checking the password. And a so-called Weaver file appears in our folder. So let's call it, for ourselves, the Weaver logic.
And let's look at how our key retrieval scheme has changed. That is, here we've added weaver_info. Well, basically it's lying in the folder under your feet, you can look at it, there's nothing interesting there. What's interesting here is forming the key for the file system. In fact, only a hash is passed to the CryptoMCU, and this crypto-processor, the crypto-chip, is a gatekeeper. So it sits there and says: aha, right, I know this hash.
So, here are your keys, go work. If the hash doesn't match the crypto-chip's expectations, it says: well, go away, you won't get the keys until you show the correct hash. So, accordingly, problems can arise here. That is, we can't offload the brute force outside the phone. So here the CryptoMCU can count the number of attempts internally. It has internal memory, so it can add up, sum these attempts. Then it has its own timers inside, it can slow down the response — which is what we run into. You enter 10 passwords, then wait a minute. And so on.
And so, theoretically, it can then tell the program: that's it, a thousand attempts, I'm clearly being cracked, I'm wiping it all. Actually, you don't need much computing power for that, it's enough to just implement a memcpy function, that is, comparing memory, matches, doesn't match, and a simple counter, and we get a whole pile of problems. There's no way we can offload this.
So, let's sum up once more what this gives us… And, essentially, why it's so hard to take apart, so hard to implement. That is, here we run into a fundamental difficulty. For us, the crypto-chip is a black box. We don't know what's inside. And the data is stored in it, the key itself. That is, it can be arbitrary, it isn't tied to the hash or to anything. Only the crypto-chip itself knows it.
Again, the whole authorization doesn't come down to getting some blob, and we can't offload it outside. An attempt counter can also appear, and errors can also be logged. That is, in fact, we can lose the data if it's handled wrong.
This is a short slide. If we just need to pull a blob out of the crypto-chip and add it to the chain, then in that case we can do something. If the password check and getting the final key happen directly inside the crypto-chip, then here we fundamentally can't do anything at the moment.
Q&A
So, basically I'm done. If there are any questions, I'm ready to hear them. Thank you very much, colleagues. Your questions by raised hand. Right, I see one in that section. Valeria Mikhailovna, I'll ask you to go there. Alexey, sit over there for now.
— Good afternoon. Let me ask a question, it's related directly to the formulas that, let's say, you presented symbolically. That is, certain operations, root of trust, for forming the key. Was this obtained empirically or from analyzing the CryptoMCU docs? This is part of the CryptoMCU documentation. Part of it is, let's say, our own learning, our research. It's a small piece of generalized information. How the keys are structured.
— In principle, this information is in open sources too. Well, maybe somewhere it's less complete, somewhere it's called differently. Since internally, let's say, within the team we formed our own names, so there's no fixed terminology anywhere here. Here a lot of people call these things whatever.
— Colleagues, any more questions? Alexey, go ahead.
— Vyacheslav, you said that's how things stand at the moment. Are any prospects expected?
— Well, most likely, if there's anything, it will all be brute-forced not on external hardware, but on the phone itself. But again, this needs separate research. Who resets, who doesn't reset the processor's state, how many times you can safely try. Maybe there will be some, but that's a completely different direction. Maybe, again, the firmware of this crypto-processor leaks somewhere, somewhere something gets clearer.
Everything moves, everything changes, the battle of shield and sword will always be there. Thank you. Colleagues, any more questions? Valeria.
— Vyacheslav, hello. My question was also exactly about key formation. You paid separate attention to whether the phone is locked or unlocked, and you said that depending, specifically on that state, the data can be destroyed if, for example, it's unlocked. Can you explain in a bit more detail what you meant? Yes, so it turns out we have a key, but let's not consider it now with the check in the crypto-chip, let's take the general case, and let's even drop the crypto-chips, it's just an auth blob, so forming the key involves the so-called phone state, that's color and is_locked.
That's the state. What is is_locked? Basically, is the bootloader unlocked or locked. Color is the phone state, that is, first it takes into account whether the bootloader is unlocked, and, I think, whether boot is swapped or not swapped, and some other files, I can't recall them right now, it's not something you hit that often. In short, it depends on the firmware state, if I generalize. That is, unlocked or locked, firmware original or not original, signed or not signed, signed with the maker's key or a different key, and all these states affect the key.
So if we changed the key, the data, let's say, if we don't reflash the phone too much, don't do anything, just accidentally installed the wrong firmware, well, there's a chance to recover the data. Roughly speaking, if we go this deep, not with the phone's own tools, but copied the data and try to work through this chain, there's a chance. But again, this is no longer the phone's own tools. So with the phone's own means, if we reflashed it or unlocked, locked it, that's it, we've lost the data for that phone.
Thank you. Colleagues, any more questions? Yes, please. Sorry, following up on the answer to the previous question, so if we, in the case where we saved the data, could we later restore it back onto the phone? Well, if we saved the data and we have the phone's unique identity, that is, the phone is in our hands, so the phone absolutely has to be there and there has to be a backup. And then, during reflashing, if something the phone doesn't like, it does a so-called wipe, meaning it erases it all. So that's a matter of luck. If we pulled the battery in time, noticed it, fine, but with live phones it's better not to do that.
— Colleagues, are there any more questions? I see a hand.
— Thank you for the talk. My question follows on from the previous one. Does it physically erase the data, or is it just that because the parameters change, the final keys change, and it can't decrypt them? Yes, that's right. At first it just changes the key, we can't decrypt them. And later on, if it actually boots up, there are internal processes in the phone, self-recovery, all sorts of other things, it's a whole world in there. And those processes may already overwrite the data. Got it, thank you.
I'd like to ask you a provocative question, as a Samsung user. Over here. What stops you from breaking the latest expensive models, the Knox folder, which, by the way, I use quite a lot? Well, models on which processors? On MediaTek, or on Kirin, or in general? Well, in general. In general? On Knox in general I can't say. What is Knox on Samsung, in fact? It's just a separate space. So there's space 0 and space 1. It's an independent space. And the problem is getting the key to that space.
So what exactly is the difficulty in getting the key to that space? Because we all remember the situation when, many years ago, the National Security Agency asked Apple to break into a phone. Apple, of course, refused. They broke it themselves. But then GrayKey appeared, out of nowhere.
— But that's a different matter. The thing is, I talk to Samsung, Samsung reps, and they claim almost 100% security for this Knox folder. Well, let's say, on devices that were on MTK and that are supported, Knox can basically be taken apart there. As for modern Exynos, I can't say. Our company basically doesn't work with Exynos.
— Alright, I wish you progress in that direction, because, as we all understand, the new models now also have AI built in, and I think they'll build AI into the phone's security too, and we all know AI can be used both for good and for bad. That's all, thank you very much. Dima, I have a question over here. Please. The question is basically about the prospect posed by phones evolving, when not only charging but also data transfer will eventually be wireless. And the absence of data transfer interfaces on mobile devices will become the basic standard for new model releases.
My question to you. Since you work more with the approach of accessing the system board and extracting information, what prospects do you see going forward for extracting data from devices that will have no data transfer interfaces at all? Apart from the fact that, the only thing left is to unlock the phone the natural way, by picking the password from other devices in the incident or the case.
And the usual photo and video capture of correspondence and recording of criminally relevant information. How do you see data extraction going forward from devices with no data interface, in short? Okay, interesting question. A bit philosophical. Let's try to, a little... Let's philosophize a bit. So, in the current situation, in any case, any device has to boot somehow. It has to be flashed somehow. All of that, I mean it's impossible at this stage, and not soon, I think, I mean it's impossible to flash memory over the air. There will still be some conductors, some circuits in there. You'll be able to get into that circuit, make changes, take a peek.
So, well, the task will get harder. Still, I think at some level there will be wires, you may have to disassemble, solder on, whatever it takes. Here I see the problem not so much in wired interfaces as in the pace of development. In general, how much attention manufacturers pay to security, how critical it's becoming.
— Well, since, let's say, our May conference up to today, as I said, for MTK in that time, what is it, 2 or 3 months, let's even say 4. Still, in that time two significant security patches came out. That's just for MTK, and that's not the most serious manufacturer. So, essentially, that pace is far scarier than wireless interfaces.
Vyacheslav, thank you very much. Let's send him off with a round of applause. Please leave the clicker on the table. Thank you very much.
Alexey Moskvichev (MKO Systems) — “Using SQLite database journals in the investigation of digital incidents”
Scheduled 11:20–11:55.
Moderator's introduction
Thank you, everyone. Thank you. And before I introduce our next speaker, let's first picture a situation that's quite familiar to everyone. You open the phone, everything's deleted, no data, of course, no records either, but we all know very well that the digital world works differently. Alexey Moskvichev, our next speaker, will tell how to work with SQLite databases. Please welcome him with applause.
Talk
Good afternoon, dear colleagues. Yes, Dmitry has introduced me.
Right, what about the clicker. Aha, it's working. Yes, today I'd like to talk about SQLite database journals. We'll briefly look at what data they store, how you can view it and find what at first glance seems to be deleted information. And we have three practical cases that practitioners in the field have shared with us. I think it'll be interesting. First of all, why SQLite? Because it's been around for over 20 years and is the most widespread, let's say, database management system for most mobile and desktop operating systems. Essentially, it doesn't require a separate server, needs almost no configuration, stores data in ordinary files, and offers fairly high performance while taking up very little space.
And if you look at it through the eyes of an examiner, it becomes obvious that a modern digital device holds dozens, sometimes hundreds of different databases, which store messages, some location data, browser history, banking app data among other things, and so on. But sometimes what's of interest is not the database itself, but its journals, which can provide equally valuable information.
So, let's imagine a simplified picture, let's call it that. A user opens some app, writes a message, sends it to someone, then edits it, and then decides to delete it. What state will the specialist or examiner see in the main database, database.db? In fact, probably the final state of the system, or no message at all, if it was deleted, but the so-called SQLite journals can preserve all three states of the system.
How WAL works: what it keeps and why to collect the journal first
Well, essentially, let's go through it briefly, I don't think I'm discovering America here, but we have a varied audience here, including lawyers, so that it's clear for them. Overall, from the point of view of the WAL mechanism, let's see how it works. So previously all user data was written to the main database file, but since SQLite 3.7 a so-called change log is kept, and all new or modified messages, including deletions, are written to this so-called WAL journal. After the so-called checkpoint procedure, they are transferred into the main database, database.db.
And that's exactly why, when examining a mobile device, in particular when obtaining a physical image or a full file system, it's advisable to analyze both the database.db file and database.wal. There's one more service file, it holds no user data, but, let's say, triaging these files lets you analyze the app data as a whole. Essentially, I've already said what records the journal may contain. These are new records not yet transferred to the database, deleted ones, some previous states of the system, temporary states, drafts perhaps. Or uncompleted transactions because, say, the device was abruptly powered off, hard-rebooted, and so on.
So that's essentially what transactions may be found there. But don't get your hopes up here. These journals, essentially, get overwritten. The lawyers might think right now that we can find and view all the deleted stuff, that it'll definitely be there. No. All these journals get overwritten. And the main goal here is to capture that journal before it collapses and before that data is transferred into the main database. And one more simple example. I text a friend at 6 pm: let's meet at 18:00, then I edit it and delete it. So, essentially, the WAL journal will keep all three states of the system, in short.
Three cases: malicious APKs from Telegram and loans taken out on the victims
Well, enough theory, I think, let's move on to practical cases. In one of the regions there was a series of similar incidents, let's say. The user who gave us this information is watching us right now. We send him a big hello. Unfortunately, he couldn't make it. So the victims filed complaints saying that loans had been taken out in their names illegally; they themselves had taken no steps to apply, nothing at all, but shortly before these incidents they had clicked on some links, and some had installed some apps.
And so, essentially, one theory put forward was that most likely there had been some kind of software on the device. To test this theory, the victims' devices were seized and sent for examination to the regional forensic center. Next, our product was used to extract the devices' file system. The methods are here. Unisoc/Spreadtrum Android and Samsung data was pulled via FFS, full file system, vulnerability 31317.
And then, in our "Malicious Objects" section, we scanned the file structure for malware. Nothing was found. And so the examiner then decided to carry out an in-depth examination of the Android operating system, so as not to look for the files themselves, which would be odd to find there anyway, but to look instead for traces of activity of those executable files.
The first incident is called "Search by Full Name bp", you'll see later. Well, the first find was the threats.db-wal journal of the Sber app, you see. Essentially, a record was found there about a file base.apk, which the built-in Sber antivirus, Sber's antivirus module, had classified as a trojan. The file itself was no longer on the device, right, but its unique identifier had survived. Note that at this stage it was already possible to draw an interim conclusion that there had been some kind of software on the device which the Sber app had detected as potentially malicious.
Next, the Android system log, which most of you are probably well familiar with, which records all information about apps installed at the current moment. There was nothing in it, no apps and no data with that identifier. For simplicity we'll call it "OK". The "OK" identifier. It wasn't there, but here's another system log of budget devices, Tecno, Infinix and itel. Note, it's palm-a, right, it recorded, it actually saved the events of the app being registered in the system on November 12, 2025, at 4:39.
Next, in the system database gass.db, that's a Google database, Google Play Protect, right, that's where information about the hash of this executable file was kept, by its identifier. And after that you could already search the databases of specialized software. They exist, you all probably know them well. Try to match it against additional episodes with similar offense types, which, essentially, is what was done. The next step for the examiner was to answer the question of what capabilities the app had, what it could do on the device. Here, in another file, Android, AndroidManifest.xml, there was information about the permissions granted to this app.
Note: sending and receiving SMS, and network access. Of course, nobody is saying that all these permissions, if granted, were actually exercised, but still the very fact of these permissions says a lot. Next, the following question that had to be answered was, essentially, how an ordinary user knew it, under what name. And here another system file of these Tecno, Infinix and itel devices let us answer that question. This app with the identifier "OKT" was displayed on the device under the name "Search by Full Name bp".
Essentially, next we needed to figure out how it got onto the device. At the very beginning I told you the victims said in their statements that some of them had followed links, some had installed something. And here the Telegram database, Pavel didn't let us down, had saved information about receiving this file "Search by Full Name bp" in one of the Telegram chats. We've blurred part of the information here. And then we needed to figure out, to find out, ideally, something about this Telegram chat, right.
This messages_v2 table itself was effectively empty, because the message had most likely been deleted, the app was removed, but in the chats table, note, the name of this chat survived, in which this APK was received: "Missing in Action in the SVO". This is probably telling, showing how, let's say, resourcefully, that's the word, scammers approach their criminal activity, they pick socially significant topics that are in demand in society right now. So, essentially, that's how the chain of events unfolded in the first incident, where it all started, remember, with the WAL journal, and then, through the operating system logs, additional information was obtained.
A similar incident is called "Photo Archive 20". In exactly the same way, the Sber app, note, kept a log, kept a record of a little file base.apk, also detected it, com.example.application. Then a log file confirmed the information that this app, com.example, had been set as the default. This is a log for the Samsung Messages app, which recorded the fact that specifically the com.example app had been set as the default app for receiving and sending SMS messages.
Next, we needed to figure out when the app had been removed. Another file confirmed the removal of the app on January 31, 2025. And then the same Google database gass.db, note, the hash was obtained. Then Kaspersky antivirus had saved information about the name of this app, "Photo Archive 20". And then again, Telegram was used as the means of delivering this executable file to the device. And here, too, the information was obtained from its database, in one of the chats.
Incident 3 differs somewhat in that the Sber app was not present on it, but the built-in antivirus also saved, the built-in antivirus log saved information about the record, and also classified it as potentially malicious. And, essentially, in addition, you see, it highlighted that this app is dangerous, there's a risk of confidential data leakage, fraudulent debiting of funds, so we recommend removing it.
And then two log files confirmed the installation of the app, note, and the removal of the app three days later.
And then another hash was confirmed in the gass.db file, note, and Telegram was likewise involved in this incident.
To sum up briefly, it all started, as I said, with the SQLite journal, in fact the SQLite journals combined with the system artifacts of the operating system made it possible to build a full timeline, to trace the chronology of events on the device, starting from the moment it got onto the device all the way to the source of this APK. So in fact it was then possible to link these events to the dates of the fraudulent actions.
Q&A
Overall, my talk is fairly short. That's it, if you have questions, go ahead. Colleagues, questions by raised hand too, if you have any. Valery, oh, Evgenia, please go ahead.
— Alexey, just a few points. One. I see you analyzed the WAL files with FQLite Carving Tool. That's some free program, it's fairly convenient to use. Does it display?
— Yes, we didn't analyze it. I said a working examiner provided us this information. The viewer used here wasn't ours, it was a third-party one. That's right. And from that follows the question of how Mobile Criminalist handles system journals. I've noticed that depending on where you launch it from, roughly speaking, Mobile Criminalist already takes them into account when displaying, say, the Telegram app, it already accounts for these system journals. But if you open it via the file browser, Mobile Criminalist shows the information as it is in the database, ignoring these journal files. Well, yes, here it's first of all a matter of taste.
If Mobile Criminalist already displays something, then of course you should probably use our viewer. If it doesn't display some of the information, then, as here, a third-party product was used. The question is more this. Does Mobile Criminalist not flag this information separately at all? Or put it this way: it doesn't flag it separately at the moment, but are there plans to implement that, so it's shown as deleted information? Actually, since we received this information from the expert, yes, of course, this task will be on the agenda.
— When exactly, I can't say. Thank you. Colleagues, more questions. Eleonora, please.
— Thank you for the talk. As I know, Mobile Criminalist can examine WAL files, not only the databases but the WAL files too, and parse them. So the problem is that when we get a copy of the data, well, like a file system copy, Mobile Criminalist examines the database, say WhatsApp or Telegram, at the default paths, but also the WAL files, and sometimes the chats end up with doubled or tripled messages. Can this problem be solved somehow going forward, so investigators don't ask why the crook sent the victim messages three times, sent some files three times in the chat? That didn't happen in reality. On the device we see it once, but in the Mobile Criminalist parse we see doubled, tripled data.
Yes, we have received such reports, yes, and our development team has them, so I think we'll work on it, yes.
— May I ask a question? Sorry, while the colleague is getting ready. The thing is, right now, besides the Anti-Fraud 3 package of bills, we have Anti-Fraud 2 already in force, Federal Law No. 210. Now a doctrine is being prepared for a system to counter ICT crime. And one of the important issues you covered in your talk is exactly countering banking fraud carried out with malware. Well, generally, yes.
And the thing is, this area now overlaps much more strongly with information security, with computer attacks in the form of malware. So there's been pushback from regulators, from the Bank of Russia, FSTEC and others, that when we turn forensic tools into a forensic system, especially under a number of departmental regulations, we need to make sure, first, that the result isn't automatic. You just showed a human analyst extracting this information, correlating it and delivering it as the result of the examination.
Are there plans to build in compliance, first of all, probably, with FSTEC, to certify some versions, maybe selectively, so they can be used in departmental systems, not as standalone tools taken out to the scene, but specifically in expert systems, and have the corresponding status: trusted software or a technical hardware-software complex. First question. And the second question relates to banking requirements and the like; one of the requirements there is that any system, whether a banking app or an analysis system, where you have to present evidence, must carry information about the absence of malware.
That is, not just get some version, say, that this is trusted software, but that the expert, doing the analysis, used, besides the forensic extraction and analysis tool, antivirus tools as well, and which ones. That is, showing that he himself didn't and couldn't have introduced any changes. Unfortunately, the current legal requirement applies directly to banking apps, but during analysis many experts get such questions in court, especially from defense lawyers, when cross-examined. How did you make sure that during your examination you yourself, or some software examining, possibly malware from that same phone, didn't corrupt your examination results?
So, two questions related to ensuring reliability. Thank you. On the first one. Did I understand correctly that as regards our product specifically, it can be accepted as evidence in court? What did you mean? Can our product be used at crime scenes?
I'll repeat it in plain Russian. First question. Will at least some versions of the product be certified as trusted software under the current infosec rules? Because the field of fighting ICT crime is overlapping more and more with general information security, banking in particular. First question. Second question. Will there be any tool-level integration with an antivirus, and will the operation logs, including the technical ones, record that the objects examined by the tool passed an antivirus scan, say, with this, this and this? Not just us writing in our report, well, nobody writes in LaTeX, in Word, we insert into the expert opinion, that we found, say, malware of such-and-such type.
But specifically a technical record that, well, is entered practically automatically or in an automated mode, so the expert doesn't need to add any extra explanatory notes, because that's already a legal requirement. Understood. Two questions. On certification, I think that's more a question for a lawyer, it's hard for me to answer, I can't say. On the second question, here I also find it hard to answer how to simplify this information, this work, for the user.
Honestly, I'm not sure. From our product's standpoint, as I said, we have a separate malicious objects section, purely the Kaspersky integration. That information is in the customer portal and in the add-ons in the customer portal that you'll be installing. What else needs to be shown or highlighted? I don't know how else to simplify it. What's here?
Colleagues, let's keep going. We have time for one more question. The young man over there has had his hand up for a while. Please, over on that side. Eleonora. Please give the mic to the young man on that side. Raise your hand higher, please, so she can see you. She's seen you, yes. Eleonora, the young man.
— Thanks for passing me the floor. Good afternoon again. My question is this: in all three cases, antivirus data came up. And it's a bit of a tangential question, just about antivirus on mobile devices. I'm curious whether it was really that easy to determine which software ultimately turned out to be malicious. Are there many false positives from built-in antivirus, and from it in the banking environment? And the second question, maybe you know, unfortunately I'm not quite an expert in this field, how strong is evidence that comes in a database format or in text form, where the antivirus logs are written, is it really reliable?
That is, can you somehow confirm such a detection really happened? Or, for instance, the antivirus logs were tampered with, and the attacker could, say, cover his tracks.
— Well, in general, data can be tampered with, you're right. But here the artifacts were examined as a whole. And so, yes, of course, this information formed the basis of the expert opinion. They were admitted as evidence in court. As for the first question about antivirus. How easy or hard it was to find all this among all that cluttered data. But here we weren't the ones who did that examination, so effectively we worked with clean data the expert provided to us. So let's check that information with him, I'll get back to you, and we'll discuss it more fully. Great, thank you.
Alexey, thank you very much. Thank you. Let's give Alexey a round of applause. And I'd like to remind everyone that right at the back of the hall is the exhibition area, and there won't be a roast today. So now we have a fairly long break until 12 o'clock, and then we'll continue. See you back in the hall.
[A break (11:55–12:30 in the program) is cut from the recording here — about 39 minutes in two pieces: before the moderator's line “And we are back…” and right after it.]
Vladimir Greshnov and Sergey Savinkov (ELETEK) — “When there are too many terabytes: from forensic images to AI analysis of multimedia”
Scheduled 12:30–13:00.
Moderator's introduction
And we're back for the second part of today's event. While the guests gather in the hall, I'll say a bit about what our next talk is about. It used to be that the main task was to get as much data as possible from various digital sources. But when you have several terabytes of various photos, videos and countless other data that today's devices store, absolutely any of them, including mobile ones now, and the big question becomes: how do you find something important, and work with all of it? And here it's going to be quite hard to get by without automation. How to move from creating forensic images to AI analysis of multimedia will be presented by our colleagues and conference partners from ELETEK, Vladimir Greshnov and Sergey Savinkov.
Guys, the floor is yours. Welcome them with applause.
Introduction: what ELETEK does
Good afternoon. We'll probably wait another minute for everyone to sit down.
— Right, now a short introduction. Our company does both integration projects and development of our own software and hardware tools. Our main areas, so to speak: we make specialized systems, hardware-software complexes, a secure execution system, software for collecting and analyzing data, tools for forensic data acquisition, operational acquisition and so on, and, accordingly, some large server-based, integration systems built to the customer's order.
Today our talk will be devoted to two of the topics listed. These are tools for forensic data acquisition and tools for analyzing heterogeneous data. For the first part I'll hand over to my colleague Vladimir, and I'll come back.
Vladimir Greshnov: the duplicator and the compact copier
— Yes, well then, greetings to everyone. As Sergey already said, today our company is represented at this event by two of its divisions. I'm going to tell you about the hardware and the software that we need for acquiring data and for working with data images.
The first system, we spent quite a long time developing it, I'd say. Last year we already presented it for the first time. But this year, or rather this month, it's going into series production and is starting to ship. This system of ours is a duplicator. A duplicator that lets us quickly and, most importantly, safely acquire data from the drives under examination.
We need it both in lab conditions, that is, you can work with drives that have been seized, and in the field, say, on some kind of site visit. This system runs off, it's powered from 220-volt mains. It can also be powered from a power bank if needed. So that's an option too, so that, say, in the field you can easily power it up and work with it. Its main purpose is copying specifically from external data storage devices, such as HDDs, SSDs and flash drives.
It's also possible to copy NVMe drives through an NVMe-to-USB adapter. Inside there's a full-fledged microcomputer. Lately we've started working very actively with modular solutions. That is, we take a certain type of brains, roughly speaking, from a microcomputer. Then we ourselves lay out boards with the ports we need, built-in hardware write blockers and so on. And we connect them together, and we end up with a full-fledged device.
As for copying. It can copy either pass-through, that is, from the drives under examination to some trusted drive of yours, or to internal storage. The internal storage is a fast NVMe disk. In terms of capacity, right now we mostly ship two-terabyte disks, but at the customer's request we can add 4-terabyte disks as well. There's no problem there, thankfully all of that is produced and actively sold now.
Our next product is our most compact copier. Basically, many of you have already heard quite a bit about it, some may have worked with it, as this product is supplied to the Interior Ministry, that's the main customer, for about the second year now.
This copier lets you copy data specifically from flash drives. Why is it limited to flash drives here? Originally it was developed as a field-operations solution, that is, officers from operational units asked for the most compact, self-contained solution you could take with you, one that would literally fit in your fist, fit in your pocket. As I said, it's self-contained, meaning it's powered by an internal 2000 mAh battery.
That's about 2 hours of copying, if you're acquiring from drives that support the 3.0 interface. If the drives are 2.0, then accordingly it will run longer.
Inside it we also have a full-fledged built-in microcomputer that makes it work. As storage here we use a microSD card, so after copying you can simply remove that card. Unfortunately, the second box isn't shown in the photo here. The box of this product, which is where that card sits. The card is removed, connected through an adapter to a computer, and there you go, and you can work with the copied data.
As for copying modes, it supports sector-by-sector copying, file copying, copying by masks, but mainly it was intended for copying files specifically. Because copying files is faster for us in some kind of operational conditions. We kept sector copying too, since a duplicator should still be able to make a sector-by-sector copy. On the next slide you can see the app for controlling both the first and the second system, though even here I'll qualify that: not so much control as pre-configuration.
On the second screenshot you can see that the operating mode is selected for the system first, that is, which mode it will copy in. Sector-by-sector, file-based, or by masks. The first screenshot is our first screen, the mode is monitored, that is, the copying itself is tracked: what's happening, what status it's in, how much space is left on internal storage. And the last, third screenshot, here we work directly with the dumps themselves or the file containers that were copied. They can either be exported. Exported to the mobile device the app is running on. Or, accordingly, you can work with that data directly on the drives they were copied to.
SATA and USB 3.0 write blockers, software systems
The next product we make, it's actually a whole group of products. These are our write blockers. Well, basically, devices everyone is very familiar with, the devices themselves are pretty simple. So here we have a SATA write blocker. We've been producing it for... quite a few years now.
Design-wise it's as simple as it gets. It can work with 2.5-inch HDDs, and 3.5-inch ones when you hook up additional power. With SSDs, basically with all devices that support the SATA interface.
And the second write blocker is our new solution. It's a USB 3.0 blocker. So it hasn't... hasn't gone on sale yet. That is, right now it's just finishing, so to speak, its final tests. Basically, it's also quite a simple device. You connect power, it powers up. It connects to the computer through port two. And through the USB Type-A port you hook up any flash device you like. If needed you can also hook up an NVMe drive through an adapter. But here, unfortunately... for now we don't support maximum NVMe speeds on this device. After all, we intended it more for flash drives.
The next group of products we make, these are no longer hardware-software systems, these are our software solutions proper. The first solution is a system we call Element-P. It lets us acquire data live from computers with Windows file systems, Windows and Linux. So how does it work? This program sits on a drive or on a flash stick. So, we connect it to the computer. The program launches, you create a task in it for the copy types you need. Say you want to grab absolutely all the files from the file system. Go ahead, you create a task, run it, and all the files just fly over to your device.
But what's interesting about it is that you don't have to grab all the files, you can search for specific files by extension, files by path, or files by MIME type. We started working with MIME types in our duplicators not that long ago. So it can detect if an attacker has stripped the extension off the file you need, or given it a different one. So the program will understand that these documents... are, say, not an audio file but a text file. And the program will grab it.
This slide shows screenshots of the interface. So, roughly what it looks like and how it works.
And our next software system is a program for working directly with data images. So with the previous solutions you can make a sector-by-sector image. And with this program you can take that image and... load it in. See what files are inside. Apply various sorting to those files by extension, by modification date, creation date and so on. And this program also supports recovering deleted information. That is, you can run a scan for any deleted files or deleted partitions.
The program can also show you that such-and-such partition used such-and-such file system, and in the deleted one, this other file system. And, accordingly, try to recover as much of the deleted data as possible. It can also work with signatures. That is, it can do signature-based search. And search by MIME types, accordingly. Well, on this slide you can see the actual screenshots of the program.
Now, this program and the previous program for copying, they're already at the final testing stage right now. So basically, in the near future you'll be able to get demo versions to try them out, accordingly, on your own setup. Well then, the block I've been talking about is finished. I'll hand over to my colleague.
Sergey Savinkov: analysing mixed data — text, audio, video
Next up is a talk on the analysis of heterogeneous data.
I'll tell you about our suite of software products, which we also integrate into systems built to the customer's spec. It has a kind of modular architecture. That is, there's a set of processors and a set of data collectors for gathering data from internet or internal portals. That could be some closed internal systems deployed on your side. Data can be combined both from the internet and from internal sources, land in a nominally closed perimeter, be analysed there, and provide the tooling.
Tooling for multi-user work by analysts. That is, these are multi-user portals with various access controls, where the analysts themselves or other staff search for information, organise it, and put together reports. Well, today's talk will mostly cover the central block, the data processing. If you're interested, we can tell you about the rest at the stand. So, the various kinds of data, roughly speaking, we split into text, audio and visual data.
We've been working with text for about 20 years. Currently we use classical algorithms for precise classification. That is, when you know in advance what you want to look for. These can be fairly abstract topics. For example, texts relating to physics, or to extremism. Anything at all. Or there can be more specialised classifiers. That is, find me specific factual mentions in the text, names, company names, addresses, assess the tone, the type of text, meaning scientific literature, free-form statements, media, and so on.
As a rule, classifiers are applied as a set. That is, specialised pipelines are built. A specific sequence of classifiers that gives the analytical result you need. That is, find me such-and-such a topic, but with a mention of, say, a particular region or country, and with a negative tone, for example, or a popular-science tone. AI tools are also now widely used for text analysis. They already let you search after the fact, that is, you first index your large sets of text data extracted from various sources. Let me repeat, this can be a single archive of data extracted from a phone, from flash drives, from computers, plus data added from open sources, internet sites.
And then you search them per a task, for example, identifying all the passages of text that mention such-and-such semantics with the required tone or behaviour, behavioural type. Audio data, what gets added here? Essentially, audio transcription gets added, a task which on the one hand is already solved these days, but on the other always needs an individual approach. That is, audio data recorded with noise, in foreign languages, with many speakers talking, with constant switching between languages, requires extra tuning and fine-tuning. We have a solution for that. You can fine-tune on your data, you can feed it your own subject dictionary.
That is, when there's specialised vocabulary, off-the-shelf transcription tools cope worse. They try to pull the transcription towards a more common word. Whereas you might have some jargon or specialised vocabulary. Accordingly, the next step is identifying speakers. For now we only do speaker separation. That is, speaker 1, speaker 2. Speaker identification is something some of our partners do. We haven't gone into that area yet. Then, accordingly, the transcribed text can be translated from foreign languages and a summary produced. That is, if you have a long recording or a set of recordings, you can get a single summary of what was discussed.
Or ask specifically, for example, whether there were any facts in this conversation or not. Visual data. On top of the first two types, an image is added. Essentially, the input is video or images. If it's video, it's split into frames and a set of tasks is run. Essentially, this is working with the frame as a whole. There's classification of what the frame shows: vehicles, people, landscapes, military, documents, and so on.
Faces are handled separately. That is, searching for specified faces, or simply clustering all the faces present in the video or images. Specialised search for specified emblems, flags, signs, certain objects, and so on. That is, you define the semantics and the frames where those semantics are present get sorted and picked out. Accordingly, the output is a large array of indexed video data, with markers placed where the objects you're looking for appear: faces, again texts, images.
And separately, we've done fairly deep work on recognising various documents. That is, these can be both printed and handwritten documents. Naturally, the quality is a bit lower for handwritten. But printed documents can also come in very different formats. That is, I'll show examples now. So, more detail on audio recordings. The specifics of our solutions. We can process recordings of any length. Streamed recordings.
There's stable memory consumption. That is, we don't try to pull the whole recording into memory at once and so on, like some solutions do. Multimodality. This means you can have a mixed recording of dialogue among several people who speak different languages at the same time. The processor picks out individual segments and, based on the language, routes them to the appropriate recogniser, the text transcriber.
Adaptability. As I mentioned, you can specify your own specialised dictionary, your domain. Which greatly improves recognition quality when you have terminology or slang. That is, you don't have to pre-train on those recordings, you just slip it a little dictionary of terms. For example, you have some radio intercepts of surveyors, with lots of terminology. And then the transcription will go very... the quality improves several times over just because of that. And now, accordingly, some examples of local analysis. That is, solutions range from a laptop up to a server cluster. For example, we have a laptop solution where processing is done in a single language.
That is, at any given moment you're processing in one language. If your recording is multilingual, processing will just take a little longer. Because it switches between languages. On a workstation you can already support multiple languages in parallel. And the speed goes up. And then, depending on the data volume, there's the server solution...
Here are visualisation examples from the portal that stores the processed archive. Well, here it's specifically about video data. But basically, the indexing and visualisation solutions are designed for all data types at once. So once you've processed all this multi-format data, you then search across all of it as in a single portal. That is, you enter queries and it searches video, audio and text data alike.
Video processing. Let me dwell on it once more. It's searching for the frames you need, classifying them, searching for faces, known or unknown. That is, it can simply process your archive, a large archive of videos, and lay out clusters of people, showing that such-and-such people appeared in it. And there they are in different, well, in different clothes, at different ages. Simply where they were, whether they were together in the same video. Beyond that, the range of analytical tasks is limited only by your imagination.
Special timestamps are created so you can jump straight to the face, emblem or statement you need. There's an editor for preparing report materials. That is, you can cut out the segments you need and make a final report video. And accompany it with a text report on another page. And another important thing is processing video not for transcription, when people are talking in it, but to describe what's happening in the video. Especially if the recordings are long or nobody is talking. Here too the range of tasks is unlimited. It's like getting a brief summary of, say, a nine-hour video where nothing happens.
From the first hour to the eighth you've got trees, wind. A fence, and five cars drive past. And from the eighth hour people come in, something's happening. So this often cuts down the analyst's working time. When they read a summary like that, they either see they need to jump to that section, or see right away there's nothing of interest. It works well, for example, on predefined topics. So when we're looking for videos on hacker topics, videos of suicides or fights. Or we're simply searching, again, when you have some street camera surveillance footage, and in 24 hours two cars drove by, at such a time, such a time, such a time, and at such a time a group of people walked past.
So, again, out of 24 hours we get three two-minute fragments that need to be analyzed in more depth.
Image classification: these can be pictures as well as, strictly speaking, video frames. Again, the groups are defined based on your tasks. They can be generic ones like portrait, landscape, maps, or some specialized ones, for example, markings on a photo, or detecting documents, printed or handwritten documents, official documents, or just documents lying on a desk, and so on.
I mentioned document recognition as a separate block because the task is so specific. We've also worked on large sets of documents, both printed and handwritten. We got pretty good results and quality. Again, it's designed for cases where, say, you... Well, no need to even mention large archives. Even the data pulled from a phone is already a huge data set, which I hope nobody's trying to go through by eye anymore, because 200 gigabytes of photos and files is just impossible to look through. Unless you have a lot of free time.
The next block is another example of the same specific task with photos, that is, for example, describing what's happening in a photo, a screenshot, an avatar, and so on.
Working with text data, analysis of large volumes of text data. Here it's shown on open data from the internet. Again, the same thing: data pulled from a phone, a flash drive, a computer. These can be full-text documents, but also chat-type documents. There's classification, author extraction, link analysis. If factual data is mentioned somewhere, that is names, organization names, addresses, phone numbers and so on. And a reporting portal that lets you save your queries on top of that and add things manually.
A huge archive of indexed texts, tens of terabytes. It lets you find the information you need in seconds, and narrow down by certain attributes. Time, as an example here. File type, where it was found, the extraction it was found in, and so on. Examples, as I mentioned, of classical text classification algorithms, which we still use today in parallel with AI. Because on some tasks, when you have a rigidly defined category, they work both faster and with better quality.
In fact, as a rule, we build classifiers for specific tasks, for the customer. We already have a large body of ready-made classifiers for standard topics. Even specialized ones we've already implemented. We have dedicated linguists for that, refining them and so on. As a rule, like I said, they work as a set.
An example of visualizing heterogeneous data. Again, a portal where data of all the types I've listed flows together. And you search through it, and it no longer matters where it came from: a closed source, some disk, or the internet. The archive lets you search everything at once. And it's multi-user. Each analyst can have their own set of classifiers. So someone works on topic 1, someone on topic 2. When they log in under their account, they see the material picked for their topic. On top of that they can refine it, add another query, narrow it by time, pick out materials for the quarterly report, the current report, the annual report and so on.
So it's a system of collections. Fairly standard modes for analytics portals by now. Accordingly, presenting the information as cards, text, maps, graphs and so on. Jumping, as in the video data example, into a specialized editor for the data type. All of that is implemented. Again, as a rule, when we're automating, say, some department, there's fine-tuning for that department's needs.
We talk about how the work is organized now, how people do it now, by hand or on some other system. And basically all these visual forms get tailored to the customer.
Q&A
That's it in brief. If there are questions, we're ready to answer.
— Colleagues, thank you very much. Your questions. Yes, I see one.
— Hello, thanks for the talk. On the first block, I'd like to clarify a point. There's a need to reproduce normal operating conditions on devices. Is there a write blocker that can fully pass through the device, the source, so it can be connected to the personal computer or laptop itself and powered on, when what's needed from the blocker is not passing through the controller's data but the data of the source drive itself, meaning the interface, serial number, SMART and so on. Because ordinary blockers just pass through the data of the controller the blocker maker builds in, the blocker's own.
Here that particular thing is partially implemented and is being implemented further in the new USB 3.0 blocker. So in the SATA blocker we don't support that. And the second question, about MIME types. Will it be possible with your product to detect, for example, archives or crypto containers disguised as some data types, images, video, some common formats, like GGUF files posing as a language model, disguised, well, like crypto containers named with a GGUF extension. Will that be possible? Yes, that will be possible. Here's how we've implemented MIME types in general. We have certain groups of MIME types that we initially set up for users.
Say, audio, archives, as you already said, images. And we also have a custom group available, where you can add all the file MIME types you need. And I'd also note here that sometimes you need the ability to search for all encrypted types possible on the device. Would it be possible to implement a search for encrypted archives disguised as some common file formats? Honestly, we haven't worked in that direction yet. So that's something to discuss with the developers specifically. But I hear you, we'll definitely raise the topic. On the second block, I'd like to clarify how this is handled.
With a local model on a rig deployed at the examiners' site? Or is this some kind of online analysis? As a rule, all these systems are deployed locally. I said we build a comprehensive solution that can even be two-tier. An open segment and a closed segment. And all of that runs directly on your side. That's why I said data from open and internal sources can be combined in one archive. Naturally, not in the cloud. This is a question about specs. What specs does the rig need to have to run the model in working conditions? And how is the context issue handled? I mean, we understand that the main problem right now is the context size, which tops out at 1 billion tokens.
How did you solve that problem? And what models are used? With what parameters? And how do you make sure that after a large data set the model doesn't start getting confused in the system prompt and in the internal instructions and so on. The hardware is chosen by the task, naturally. That is, how much you want to process and how fast.
I gave an example from a laptop on up. As for all the other questions, how should I put it. Again, the model searches the whole archive. So if you have a terabyte, it's already indexed. It has a schema of that archive. And it searches within the context depth, that is, whatever you ask it. But it's also limited by the standard things. So as a rule, it's, for example, two GPUs that run a model of 38 or 72 in size. So there's nothing special going on there. So the query context is, naturally, limited. But for analyzing an archive that's already structured, with a schema built for it, that's quite enough. So it performs the search across the whole archive.
Colleagues, any more questions? Let's do it by raised hand.
Could you tell me, please, about the USB blocker. What's the technical solution? Is it a microcomputer in there too, or a signal processor? Look, we use a microcomputer there. And on top of that, our modular solution. We don't use the computer itself, just the brain of it. Then our own board is laid out, which connects to it. And that's how it works. So the blocking is still done in software? Inside the device, yes. Thank you.
— I see a hand, coming over. Can you keep it up? Because I lose it among the rows. Thank you.
— Hello. A question about the Element VD software product. Can you name its key advantages compared to competitors? Or is it simply an alternative because it's Russian software? Well, first, yes, what you said. It's specifically Russian software. And probably the main advantage it has is that it runs on Astra Linux. And working with, well, I'm not sure now, R-Studio. It seems to work with E01.
That was our main, large-scale task. Working with E01 files on computers running the Astra Linux operating system. Got it, thank you.
— Thank you. Right, colleagues, we have time for one more question. Slava, I see you. Alright, fine, two. Okay, colleague, you go first. You raised first. Yes, thanks for the talk. I wanted to ask, when you came to visit us, I already asked you then how you handle interruptions while making an E01 image. Back then you said you weren't aware of such solutions. I pointed out that they're already on the market, and you wanted to look into adding them. So if we're creating an E01 image and the drive drops out, everything stops, with you it all had to be started over, but there are products that can resume.
Well, if you tell us which product that is, we'll at least be able to see roughly how it works, because, based on the information we have, the E01 format itself simply doesn't support interruptions, since checksums are calculated within the format itself, and after an interruption that can't be done right. I already told you: Belkasoft, Tableau, you can look them up. Does Tableau really have that? Yes.
We'll look into it, alright.
— Right, good afternoon. Tell me, please, how is the emotional state assessed? You had that written on your slide. Well, the emotional state by the tone. So you mean audio recordings? No, no, you had it for audio. For audio recordings. Well, by timbre, naturally. By timbre, pauses. Well, some people just speak louder, some quieter, and so on. No, but it does assess it. Was there scientific work, or just a guess?
Scientific work was done, there are already standard solutions for this. Well, I see, okay. And how will it be described, in what words, a fight, say? Waving arms, or something else, because someone standing there doing something, is that enough? Well, it assesses the dynamics of the video, so if, roughly speaking, it describes that a fight is taking place, because... Well, no, it could be hitting, or something else, you see? No, well, it doesn't break down the type of fight inside, it just says a group of people... That's not what I mean. What's it for? Say, a large number of video images that may contain some unlawful actions.
And if it outputs that for you as text, then you can search by that text too. Yes, yes. But what do you search for? Well... Exactly. What words does it write it in, so it's clear... Look, as a rule, as a rule, I had an example there, a query isn't written as a single word, but as a group of words with, well, special linguistic terms. And when an analysis portal like this is set up, a query library is built.
Because the words written can vary. It may just be people moving actively, approaching each other, right, so in fact the word "strike" may not appear explicitly. "Fight," because it may not be there, they may be standing close to each other, shoving, and someone just stabbed somebody with a knife, for example. And that may be written out explicitly. But on top there's the linguistic query, which must account for all these nuances, like, well, for example, worked out on a test sample, right. Well, so first you need to study what you output, then think about what... Well, not what we output, but what we want to search for.
And then the linguists themselves tune the queries for that. So it won't be just one word or two, it'll be something like a mini-classifier, in effect. No, I understand, it's just that usually linguists don't do this, but computer experts. No, it's better with linguists, because they better understand the semantics of the operators, how to link the words. And one more thing, in what form is the dictionary built, for specific terms? Is it just text? Or do you still have to pronounce the word plus text? No, no, no, it's text.
In these kinds of audio recordings you can run into specific words like that, right? I mean, certain terms, slang, right? Because with standard transcription it may just not have that specialized word in its context. It's more likely to recognize it as some more common word. When you specify that there are nicknames, or, I don't know, special terms, it will prioritize looking for those, yes. Yes, you just enter a word. A word, Enter, a word, a phrase.
— The closing one, sorry, that's ours.
— Testing, testing. A follow-up to the previous question. I'd like to know how you've implemented reproducibility of results. A common problem with AI-based analysis is that the output results change. How have you solved that? Obviously, you can lower the temperature, the model's tendency to make things up. How have you implemented reproducibility? So that, for example, a video is first analyzed one way and you get certain results, and then sometimes the model starts giving different results after analysis. And how have you implemented that? This is very important for expert examination and analysis results. Processing the same video twice, or different videos?
Yes, yes, yes. The parameter there is hard-set. So the reproducibility will be very high. Overall. It's like this making things up, the imagination of the neural net. With this kind of analysis it's practically... Well, it's not switched off, you can't switch it off, right? But it's heavily restricted. Because otherwise, when transcribing audio or video, for example, you'd get arbitrary text resembling something. So we restrict the parameters strictly there. Well, we do preprocessing, slicing. I think, for audio tracks, right? And accordingly, its flights of fancy are very heavily limited, yes. So each language has its own set of models. For example, for slicing and preprocessing a video file there are separate small models.
So there's a whole stack of models.
— Colleagues, thank you very much. The guys have a booth here. And you can talk to them all day today. Let's send them off with applause. You can leave it there, please. Thank you.
Dmitry Boroshchuk (Beholderishere.Consulting) — “Media forensics: improving photo, video and sound”
Scheduled 13:05–13:35.
Moderator's introduction
And I'd like to introduce the next speaker. But first, a short preamble. Sometimes we come across various digital, and not only, media files. That's more accurate, knowing what the next talk is about. But they come in completely different quality. And how to work with all that, our dear friend Dmitry Boroshchuk will tell us today. Please welcome him.
To start: what the room uses to clean up photo, video and sound
Hi everyone.
Seriously. Let's start with a chat. How often do you come across a video or audio file where everything is bad? Raise your hands. One. Oh, great. And what do you clean it up with? Colleague, I feel your pain. What do you clean it with? Right from your seat.
— Well, give him a microphone. Yes, let me run around. I could use a second person. Well, usually we apply a model that improves quality. It's, um, minimally, I forgot what it's called. It increases the resolution. And often we use neighboring frames in the video to figure out what's visible in one part and in another. And so, using natural intelligence, you get to, say, the license plate you're looking for. You mentioned the AI, but do you file that in the case too? Like, the AI said: here's the plate number. Well, of course not.
— Good. And what about sound? Over here, another colleague. There was a hand. Dmitry Vladimirovich, I saw another hand there about photos. It was raised high. Actually, this question came out of those ordinary conversations that start in the kitchen and begin like this. Well, we pulled the camera footage, took a look, nothing there, and moved on. Yes, colleague. Originally about photos, I wanted to say... With Amped FIVE and other software solutions, but as for sound... Which ones? Well, various algorithms, for example Richardson-Lucy... Photo Expert, yes, great tool.
— As for sound, I enhanced it exclusively with a neural network. Well, in this case, as for enhancing the sound itself with non-AI methods, there's a compressor... Reverb and so on. Well, that's for vocal material specifically, so to speak. Vocal-instrumental, practically. Well, something like that.
— Okay, so, who else has what problems? And what's your problem, colleague? Well, for sound it's Audition, iZotope by noise pattern, if we have a lot of noise. Naturally, we pull it out, or otherwise we just keep listening. I don't use the AI, because it distorts things heavily, and even the transcription is complete garbage. Yeah, I agree. And what's the biggest problem you run into? Can't see anything, too dark, or too bright, noisy, too little... Poor visibility, right?
Look, today I suggest we talk about... Why did I ask about the software? Because most often, at the moment we need it, the necessary... Licensed software is what we don't have. Or it's just not on hand at the right moment, right? Or we simply don't have access to it. And mostly the detectives wrap up the whole investigation like that... Well, can't see anything.
Free tools only, and no AI
And today I suggest we talk about the software products that are freely available to you, openly available. You can download them freely, without cracking anything. We'll only talk about free tools and only about widely available ones. And, essentially, so that all our actions are, as forensics requires, reproducible afterwards. Why can't we use AI? Yes, the AI is great. And most likely, using AI looks pretty much like it does in that TV show.
It's a magic button. But unfortunately, there is no such magic button. And more often than not we have to be a bit of a photo editor, a video editor, and a bit of an audio specialist, to clean up the moment we need in a photo, in a video, or a recording, say, from a voice recorder. We'll be talking strictly about algorithms. The slides will have step-by-step instructions on what to do for cleanup. And this part will follow the questions that field officers usually come to us with.
Video: pick the codec, check the metadata, rescue a dark recording
First question: we got a video, and a standard player won't read it. That happens a lot. It was pulled off some DVR or some specific piece of recording hardware, and a regular video player won't read it. What can we use here? Of course, we need to find the right codec. I'm going to play Captain Obvious a bit here, but without this foundation we won't understand what comes next. Two excellent, long-established video codec packs. K-Lite Codec Pack for Windows, and a universal player that chews through practically 95% of all media content, called VLC Media Player.
If we still couldn't play it back. Then we need to transcode it. Because the file may be something specific, with a specific extension. And the signatures won't directly tell you what to play it with. There's an excellent solution called Shutter Encoder. It's completely free, and it's very convenient. And its main feature is that it lets you transcode any video into any video. Without losing quality. You surely know that a lot of data can get lost during transcoding. And Shutter Encoder is exactly what will let you, at minimum, transcode without quality loss. At most, see some metadata associated with the file that we can then use in our further work.
MediaInfo is another program. Which, as you've guessed from the name, lets us pull out every possible bit of metadata. To check, for example, whether our file was modified on its way from the source to us. When it was made, what it was made with. And lay all that out visually. The next issue that comes up: the video needs to be made readable. But you can't apply lossy compression. So as not to destroy the fine details. We go back to Shutter Encoder, which we talked about earlier. Next. The defense submits a video from a phone. We need to check exactly which phone it was shot on. Another metadata viewer and extractor is Metadata++.
Also a free application. You've surely heard of it. It lets you pull metadata out of practically any file. And pull out as much of it as possible. And thanks to it, you can extract metadata not only from media files. But basically from everything else too. And it's a pretty decent cataloger.
Next problem: the DVR footage was recorded at night. The image is very dark. You can barely see anything. How do you brighten the scene without quality loss? There can be a lot of steps. The first thing that comes to mind is to fiddle with brightness and contrast. And somehow try to brighten it. In 99% of cases that will get you nowhere. Here you need to understand how the picture is formed. First of all, any photo or video image is a set of pixels. If there's no data in the pixels, then you won't pull anything out. So here we'll be working with colors. With the colors that may be too overexposed or, conversely, too dark.
Another free program. Who here does video editing? The young ones. I've seen a lot of young colleagues. It's called DaVinci Resolve. An excellent video editor. TikTok clips are made with nothing else. But on top of everything else, it's also an excellent tool for working on the picture. Next we'll have an algorithm. I won't dwell on it too much. You can try it out yourselves, after all. So we can squeeze into the time slot.
Next point: license plates. Oh, look at all the phones going up. Colleagues, I'll share a link afterwards. This presentation will be much easier for you than watching these slides with my mug on stage later. So, reading a plate number. How? Actually, there are a lot of additional factors we have to take into account regarding how that plate ended up in the image and what was happening on Earth at that moment. The most common thing you run into is blur. Motion blur because the camera was moving. Someone was shooting with shaky hands or on the run. Or the car was moving. Here we'll need to sync up with the picture. Take the neighboring frames and, essentially, try to reconstruct that plate by removing the blur.
On the screen, as an example, there's a program called VideoCleaner. A fairly old, old software product, a forensic one, a free software product, which is probably about 7 years old now. But it still lets you perform some simple operations. You can also do this with that same DaVinci Resolve. And the Mobile Criminalist folks are supposed to release an e-zine tomorrow, a sort of electronic magazine as a little website, where there'll be my article on how to deal with these kinds of things in more detail.
So, VideoCleaner is a great help. DaVinci Resolve is a great thing. Moving on.
A pixelated face, a licence plate, a reflection in glass
In the surveillance camera footage the perpetrator's face is a pixelated blob. Can it be made recognizable for identification? Can it?
Now, I won't talk to you about artificial intelligence here, because most of the time, when we have a 64 by 64 pixel image, at best the artificial intelligence starts fantasizing. You get all sorts of interesting things, but unfortunately they have no connection to reality whatsoever. So we need to try to pull some additional data out of the neighboring pixels. It doesn't always work, but using, again, a fairly old super-resolution method, you can try to do interpolation and at least bring out specific areas that will help us identify the person. Moving on. There's a glare on a window in the frame.
Feels like the show "Sled," right? Which possibly contains the criminal's reflected face or contains some reflection that will let us identify, well, at least the place where this is happening. Actually, yes, a reflection contains a wealth of information, which, unfortunately, for some reason colleagues don't notice. And here, first and foremost, we act as... photo editors, where our main task is not to improve the picture, but, on the contrary, to degrade it so as to pull out that very reflection from where you'd think it couldn't be. So pay attention to those reflective surfaces which, hypothetically, might contain those files.
Just as free and wonderful: GIMP. You've surely heard of it. A great alternative to Photoshop. A free alternative to Photoshop. And the same DaVinci Resolve, so we can pull together, at the very least, the frames with a convenient, or rather a clearer reflection. Next point. Several witnesses filmed a fight on phones from different spots, and we need to somehow reconstruct the events. There are lots of small software products, free, open-source products, that let us play several video files in sync and tie them to the location. But again, so you don't spend ages hunting for them, use that same DaVinci Resolve and lay everything out on the editing tracks, on the timeline tracks, so that later you can sync up on some event, a flash of light or some loud sound, and switch between those tracks during playback to see what you can piece together from the other camera.
Sound: what can and cannot be recovered; filters in Audacity
Well, let's move on to audio. What problems do we usually have with audio? Nobody has any problems?
— Never had any.
— Noise? Right, very quiet, lots of people talking, everyone at once. And, again, something has to be done about it. And let's start with the fact that first we need to understand what we can pull out and what we can't. We'll remove sound that completely masks the speech. There's no way around that. We can't paint in words that are missing. The AI, unfortunately, doesn't work here. You can play around with processing, but we're unlikely to get a reproducible result. Because every time we process it with that same AI, we will, unfortunately, get completely different versions. But we can even out the volume, improve intelligibility. We can, after all, remove those moments when the mic rubs against a jacket lapel.
The most common artifact that gets in the way of recovery. And we'll do all of this purely based on the physics of sound. Let's first understand what can get in our way when playing back some speech. It could be some constant background. As my colleague said, we take... Just a noise sample, remove it, and listen. Yes, that's an option. When, for example, we have a loud air conditioner running constantly in the background, or some monotonous noise can be heard.
It could be some low-frequency rumble. For example, when we try to record something outdoors. When the urban environment adds the roar of cars, a construction site working nearby. It's wind. Actually, wind and the noise from constant rubbing against some surface are roughly the same. It's hiss, it's rain, it's city noise, it's clicks, and it's clipping. When the sound level exceeds everything and our recording equipment simply cuts off the frequencies. Accordingly, we can't get it back. As I said, we use Audacity, a free audio editor, which lets us remove all this interference, all these noises. Again, each has its own filter here, where we can try playing around with removing it.
So, let's try to answer the question, when some audio comes to us. What kind of noise are we dealing with? Is the noise constant or does it change? What is it mostly like? Clicks, hiss? Where exactly on the spectrum is the main interference concentrated? Does it mask the speech completely? For this we go down the path, essentially, of needing to visualize everything we see. We turn on two types of spectrogram, which will show us exactly where to look for those noises, where to look for the voice we can isolate.
Well, and then. Usually the noise is monotonous, it hardly changes, clearly audible in the pauses, the spectral picture is stable. Here everything is fairly simple. As my colleague said, we take Audacity, we cut, we look for a spot where there's no useful signal, we cut out that noise, feed it to the automatic noise remover, and it removes that monotonous noise from the recording, again, under ideal conditions. It rarely happens that the noise is monotonous and, in principle, all these steps are enough to get it done. Next is rumble. The noise feels like vibration, rumble.
When the main energy, the sound energy, is at the bottom of the spectrogram, when the recording is overloaded, when the mic gets touched. A different story. We go into effects, we go into the equalizer, and we try playing with a high-pass filter to remove that very noise. With wind it's roughly the same. When there's wind, because wind and rumble, in principle, give us the same picture.
Next point, when the noise sounds like hiss. Here it's the upper frequency range, which is on the spectrogram, where we can also try to clean it out with the reverse noise reduction for the low...
through a low-pass filter. City noise. The noise changes constantly, sources appear and disappear. The spectral picture is ambiguous. Here we're already working across the whole spectrum. We take the equalizer and try, first of all, to mark the places where our noise is especially pronounced. Which type of noise is especially pronounced. And we try to work around it. Using a graphic equalizer. And, as the final thing, when we need to bring volume swings to one constant level, that process is called compression, which you can also do thanks to this algorithm. And normalize all the audio you had in the thick of it.
That's actually everything I wanted to tell you today. If we happen not to know each other, I have a little hobby, two Telegram channels.
Q&A
And on that channel, most likely, this presentation will appear in a few minutes. Questions? Before the questions, I'll point out that the article you wrote, Dmitry, Vladimirovich, is already up on the portal with the QR code we showed today. So it's already available, basically. And now, colleagues, yes, questions please. I'm looking at your hands. It's as if everyone... No, there are hands. That's usually when either everything's clear or nothing is.
— Good morning. I have a question, probably a rather silly one. If we have a recording where nothing can be made out, and it doesn't matter if it's audio or video? There's an investigation into some case. That recording was cleaned up and submitted as evidence. Can the defense lawyer object and say the recording was altered, or not? I mean, will it be admitted in court as evidence? Or is that a question for the lawyers? If you used the AI, then the lawyer will say: "Yeah, some weird crap, and it made it all up."
Since all the operations we've just been talking about are reproducible, then after the expert testifies... After the expert or specialist testifies, the court usually accepts it.
— Dmitry, thank you for the talk. You're welcome. As a former forensic audio examiner, I can just say that what you've presented is roughly the level of maybe 2006 or 2007. I wish you to pick up more methods. In particular, if our noise has structured harmonic components, they're very easily removed without damaging the speech. There's a little program, sadly the author is dead, called Justiphone. Yes, a wonderful thing. It removes harmonic components beautifully. There's a demo example of a conversation against an organ background. When we have a clearly defined interference, we can subtract it without damaging what lies underneath.
Second, on video. Here too I'd advise colleagues not to trust too much what's visible at first glance, because you didn't cover all methods. Absolutely. And on the practical side, we had to reconstruct a shootout. You probably know. There was a shootout at Wildberries back in the day. And there were about 45 video recordings from different sources: fixed cameras, dashcams, Dozor body cams, mobile devices. And we had to manually stitch the video streams into a wider frame, so that afterwards, instead of switching, we could watch them at once; of course it was hard on the eyes, but we analyzed who was shooting from which side and whose pocket that pistol had been in.
Here's what I want to add. Just an addition to your example. It's simply very good when we don't just switch between video streams, but view them simultaneously, in a sort of multiplexed mode. Thank you. Yes, thank you very much, first of all, colleague. Yes, 2006 to 2008, absolutely, the same VideoCleaner. The question is different. What can we use when we have nothing at hand? And what can we use when we don't really have any special expertise? But it has to be done. One moment. About event reconstruction. One moment. Dmitry Vladimirovich, we're still being filmed. May I? Oh, sorry. On stage, please. Yes. About event reconstruction.
It really is a problem when we have several sources, we have close to a hundred witnesses we've collected video data from, which then has to be placed on a map. And also with elevation, and the viewing angle, and tied to the timeline. A little teaser. My team is right now preparing exactly such a tool for reconstructing events from video data with geolocation to the terrain. It will most likely be free, and you'll be able to use it. Something like that.
Dmitry Vladimirovich, thank you very much. Unfortunately, we're out of time. Let's see Dmitry off with applause.
And now I probably have some not-so-good news for those who are with us online, because for today our event is over. But only for those who are online. The next parts won't be streamed, and we'll be back with you tomorrow at 11 in the morning. But for those here in the hall today, we're moving on to our, let's say, more closed-door, more hands-on part of the event.
[End of the day 1 stream. What follows was not streamed: the talk by Ekaterina Mingaraeva (SEUSLAB, 13:40–14:15 in the program), the closed-door part of day 1 (after 14:20) and the opening of the 2021 time capsule. The next line is already the opening of day 2, 4 September.]
Day 2 — Friday, 4 September 2026
Day 2 opening
Dmitry Yankovoy — moderator, and Valeria Vakhrushina (MKO Systems).
Housekeeping and the venue
So, dear friends, good afternoon everyone.
[Between the greeting and what follows, the local recording has 25 minutes of technical pause (the moderator goes on to apologise for the trouble); on YouTube it is cut.]
Please excuse the minor technical hiccups, but we're ready to begin. Also, hello to everyone watching us online today. I'll run through a few technical points, but first let me introduce myself. My name is still Dmitry Yankovoy. Today I'll be moderating the second day of our wonderful conference. As I mentioned before, this is the talks area. A bit further on there'll always be a lovely coffee break for you: tea, coffee, snacks and so on. A bit further still is the booth area itself, ours and our partners'. Yesterday there was a very funny moment, because I kept sending everyone there, but people kept going to the registration girls.
Today you can't miss it, the bright yellow booth is right behind the guys. You can go up to them, they'll tell you everything, and we still have gifts. Besides that, we have two exits in operation. Or rather, one is for entering and the other for exiting. So if there are smokers, please go out. Go out to the left, then walk around and come back up on the right. I won't drag this out, because we've already run a bit over. And to open, I invite our marketing director, Valeria Mikhailovna Vakhrushina. Let's give her a round of applause.
Welcome — Valeria Vakhrushina
Good afternoon. Dear guests, we're glad to see you on the second day of our tenth anniversary Moscow Forensics Day conference. I hope yesterday was interesting. For those who weren't here, glad to see you. Well, I suggest we get started, and we'll chat later. Have a good day, everyone, and enjoy the talks.
Igor Bederov (Internet-Rozysk, Cybersystema Group) — “Ad intelligence and other ways to find someone by their phone”
Scheduled 11:05–11:35.
Moderator's introduction
Thank you, Valeria. I won't drag things out much either. I'll just introduce the person who's about to speak. Every time we pick a talk topic with him, it's always something new, interesting and unusual. Every time he sends me a topic, I even find myself wondering: "Damn, what will he come up with next time?" Please welcome with applause: Igor Bederov.
Talk
Hello, colleagues. Thanks very much for coming. Glad to see you all at yet another event, an anniversary one this time. How many times now? My fourth time taking part, right? The fourth one, right? Great. So what's on today? In previous years we've talked about Telegram users, Telegram channels, and investigating websites. And indeed, when Dmitry and I were thinking over the topics, we decided, why not take a topic that's partly related to the Mobile Criminalist product line. And we picked a topic that shows us how, using competitive intelligence methods, using OSINT methods, one could track a mobile phone's movements.
Again, using OSINT. I'm not law enforcement, and I'm in no way encroaching on that delicate, mostly unnecessary domain. I'm talking about how you could track mobile phones using open sources of information. Well, here's a brief bio. You can take a photo or not. You'll find me online anyway. First, let's go back a little into the past. And in the past, up until about mid-2018, everything was quite interesting, legal, simple, elementary even, I'd say. Because back then, tracking mobile phones took no effort whatsoever. Let's recall.
How tracking worked before 2018: SMS centers and HLR lookups
Before 2018, well, what was it, 10 years ago, 15 years ago, we had SMS centers. In the SMS centers you could get the LAC and CID, that is, data on the base station's coordinates, practically in real time. So you had to make a few requests on the phone number, a ping SMS, then an HLR lookup, and you got the exact coordinates. A bit later the situation changed somewhat. The mobile operators started returning some kind of code. Instead of the LAC and CID data. But that turned out to be a geographic code too. And you could drive around major cities for a while, collecting codes that were tied to a specific area.
And then, as you got those codes via HLR lookups, you'd tie them to one territory or another as well. But in the end, in April 2018, that Overton window closed. For better or for worse. It survived another year in Ukraine. And, say, in the United States. But gradually it faded away. And the SMS centers, instead of LAC and CID codes, or even unique codes tied to map locations, started returning just their own SMS center's data. Which naturally was no help at all to us in our work. But at the time it certainly had a wow effect. Because back then, we remember countless private detectives advertising "flash" services, call detail records.
They sold all that to their clients. For 10 or 15 thousand rubles, some for 20. Those are all the old days. We did it for 29 kopecks. We got that kind of information. And we didn't need any corrupt connections whatsoever. We got it completely legally, through an equally legal SMS center. Then the question came up. How do we keep getting roughly this kind of information? How do we keep tracking mobile phones? And we concluded: with OSINT methods, the ones we're talking about today.
With access to the phone: parental control and social engineering
You can do it using several approaches. And these approaches split into two categories. Category one: those requiring access to the device itself that we're tracking. And category two: those that don't require it. Let's go through them one by one. If we have access to the mobile phone. Well, we got it one way or another. Got physical access. So what can we install on that phone? We can install parental control software. The simplest and most popular option. Kaspersky, Google's standard app and the like. There are lots of them. They all let you track a child's or an employee's device quite well.
Finally, there are mobile operator services you can subscribe to. If, for example, you monitor your employees, mobile operators can well give you the ability to track, to track their movements, to track relatives' movements. And that's also quite smart, effective and simple. And finally, we have the live location sharing feature. Here we use most of the services we have on our phones. Let's take the simplest example. The Telegram messenger, being blocked in the Russian Federation. You can open Telegram on the target's phone. Start a chat with yourself. Set up continuous sharing of their location to your own phone.
And then delete that chat on the target's phone. Meanwhile the sharing to your device will remain, it persists. And you'll be able to keep monitoring that device's movements. Well, at least until they turn off GPS or reboot the phone. That's as far as getting physical access goes. Here, again, there are all sorts of nuances to getting it.
Of course, there are very experienced people who have Mobile Criminalist. Maybe other software products you use for unlocking. We, of course, don't have that. Even though I used Mobile Criminalist for a while. I love and adore it. But in the field we often can't work with it at all. And even if we're brought in somewhere to give practical help to the police, we work with what we've got. And what have we got? We've got social engineering to unlock the phone. And it tells us that in most cases a person is willing to share their phone password. The "can I borrow your phone for a call" methods always work.
You can always try to watch them enter the password. There's a certain repetitiveness in those passwords, in pattern-lock inputs. There are all kinds of powders, baby talc. There was even a really funny case shown on TV once, where a girl was given Turkish delight to eat and then handed her phone. She tapped around, and naturally the delight left greasy marks on the phone in the spots she pressed to unlock it. And so on. Methods like that are described on Habr in plenty. And you surely know them without me. So let's move on to the next direction.
Without access: digital traces, logging, leaked data
What do you do if you have no physical access to the phone at all? And here too. OSINT says that surveillance is, in principle, possible. Maybe not of the phone itself, but of the whole set of devices and information associated with its owner. Still, what can we do here? First, the obvious social engineering. We have a phone number. We can call that number. And we can ask a question under the cover story of a courier, a flower delivery and the like. Ask when they'll be home. It's something, at least, a way to track. Rough movements or presence at a particular place can be established here.
Then we have a huge number of online services, which we also identify by the mobile phone number. Social networks, Google and Apple accounts, various messengers, and so on. All that social activity which lets us find that golden grain in the huge amount of information a user publishes every day. And that grain can be quite useful. And not all of it is clear to the user. Because something we upload to Google's cloud, almost without noticing, turns out to be tied to Maps. Our numerous reviews, comments. The negativity we dump on, say, shops. And all this information taken together.
And we do a lot of vetting, for example, of people being hired into positions of responsibility; we assess them. Because we pull out their entire digital background, including subscriptions, public messages in chats, channels, communities, walls and so on. And then we analyze it for, say, loyalty. Here we move, let's say, from loyalty to tracking movements, linked addresses, persons and so on. And again we use a neural network on all content the user left on the web. And that lets us build a map of their movements. Their heat map, their connections, ties to particular addresses.
We never fully realize what volume of information we simply spill about ourselves online. And that even goes for some criminals. The second thing that comes up here is the use of various logging tools. Logging means getting a user's digital fingerprint by virtue of them visiting a particular web resource. Logging today makes it possible to obtain not only the IP address and data about the device itself, its browser, the screen resolution, language settings, Java, cookies, Flash files and so on, but in some cases, with extra settings, in the interests of law enforcement, it naturally lets you see the accounts tied to the log, to the device and logged in within the browser, for instance.
They let you see geolocation, in some cases get a photo from the front camera, a voice sample and similar data. But what interests us here is geolocation, so organizing mass messaging via bots, via ad networks, via email campaigns containing logged objects, lets you, to some extent, track a user's movements, their presence at a certain time in a certain area. But this is still, so to speak, a one-sided game. This technology runs on the capabilities of HTML5, the Geolocation API, which is actually already 12 years old. And it lets you fix our location using GPS, Wi-Fi, LBS from cell towers, cellular networks, and finally the IP address.
It does require permission, of course, in most browsers, but on mobile phones that permission is often on by default, which makes it possible to track where they might be. Then, another interesting idea within OSINT research was that, given the huge number of leaked passwords, leaked emails linked to those passwords, it became possible to use the search systems for a lost phone that most operating systems offer us. What do we mostly have? We have iOS from Apple and Android from Google. Both systems let us get access to a lost phone, to its location, if we know two parameters.
The login, that is, the email used to register the account, and the current password. There can't be any two-factor authentication here, simply because there's nowhere to confirm it; your phone is lost. So, knowing those two parameters, the email and the current password, you can get information on where that phone is, provided, of course, that option is enabled in your device's settings. Find my lost phone. In that case, yes, you can quite easily see the phone's location in a given area and track its movements online, which is pretty convenient. I, for one, keep an eye on my daughter that way, bypassing the parental apps that annoy her terribly.
Radio interception, Wi-Fi radars and AirTag trackers
Next point. Every one of our phones and all our smart gadgets constantly broadcast a huge number of signals outward. Just imagine, today we're all hung with phones, smartwatches, smart earbuds, smart rings by now, headsets and things like that. All of these broadcast Bluetooth identifiers, LoRa, MAC addresses, floating around. And all that can be seen by an outside observer. In other words, that outside observer may have a device running, say, WiGLE, nRF, LightBlue, BLE Radar, or other software that detects networks and radio broadcasts from devices like these around them.
And they really can, to some extent, track your movements. Well, the simplest example. Surveillance is under way. Someone follows the target. And to make sure the target doesn't get far away, or to check they're at the address, they can turn on WiGLE and check for a particular device with a particular MAC address in the room, within range of their phone. If we have a grid of such devices, like what they proposed installing in shopping malls a while back, Wi-Fi radars, then using that grid of devices, collecting centralized information from them, we'll be doing the same thing mobile operators do when they track mobile devices moving between their cell towers.
Here we'll see the whole spectrum of the Internet of Things, IoT objects that may move between shopping malls, and, accordingly, record that information. And this is no longer a phone number that can change, this is actual detection of the device by its MAC address, its modem's address.
So we're basically done with that. What else comes up here in terms of interesting elements around surveillance? Over the past year I've had several cases involving the use of trackers, most often AirTags, used to spy on individuals, for a possible attempt, as the victims there claim, on their lives, and other nasty things like that. What's interesting and what's the problem? First, most trackers today are detectable. Modern operating systems, iOS and Android from version 15, detect most trackers around them. If they don't, for example, Apple is tuned to its own maker's trackers, then you can install one piece of software or another on your phone that lets you detect trackers of this kind.
Those software tools are listed here as well. So detecting a tracker is no trouble at all. But what happens when we detect it? We inevitably come to believe we're being followed. Meanwhile, if we take other Apple IoT objects, say, earbuds from that same company, which in exactly the same way, using the device network of that maker, can reveal their location, then here, let's say, there'll be far less suspicion of ongoing surveillance. Finding a lost earbud in no way leads a person to conclude that they're being followed. So it may make sense to use for surveillance not only classic trackers, but also the IoT capabilities that modern IT companies provide us.
And finally, among the things that matter today. We all, of course, understand how the numerous IT corporations are watching us. And marketers, and everyone who today collects without any control our digital fingerprint, collects data on our movements, our interests, information about our purchases, our taxi rides, and scrapes our app data. All that is certainly evil, on one hand. On the other, they make our world a bit more convenient, in their own sense of convenience.
The flip side of this is that the advertising identifier assigned to us in the course of this tracking can also be used by private individuals to carry out that very surveillance on us. So we're no longer just hostages of some IT corporation that, yes, may decide to hand over data on our movements, as Google has done repeatedly, to the police. Now we also have plain private individuals who use a particular way of getting advertising identifier data. And that data can be used to track movements.
ADINT: intelligence from advertising identifiers
This methodology is called ADINT: advertising identifier tracking, intelligence from advertising sources. What does it let you do? It emerged in late 2017 in the depths of the Paul Allen School of Computer Science at the University of Washington, it was very quickly adopted, or picked up, let's say, in our research community. And back in 2019 I had already written my first paper on it. What does it let you do?
Researchers from Washington asked themselves a very good question back in 2017. Say I'm a mobile user. I went to some website. On that website I was shown a banner ad. Then I went to a different site. But that ad started following me across all the other sites. So there's probably some kind of identifier that somehow lets all these external sites pass information about me and my location between them. And in their research they proved that not only users' location and their movements are under control. What's under control is collecting geolocation data about the whole population of users.
Establishing their gender, age, interests, income level, relatives. Establishing connections between users. That is, everything we used to associate with the results of serious analytics on telecom operator data. Today a machine does it all. And ADINT gradually started to take on a more or less definite shape. The first software products appeared, in Israel naturally, that let you use the advertising identifier to set up surveillance. Of course, the data there was enriched with various OSINT methods, and with plain old data leaks, to raise efficiency, to improve how much data gets collected and displayed from users about the target.
In fact, what goes into ADINT as input is even, and this is the key point, we're used to our attacker very often changing phones, changing email addresses, changing phone numbers. Given the position of the IT corporations, they'll ignore all that and keep linking our digital fingerprint, stitching our digital portrait, aiming to ensure that, despite us swapping out all our identifiers, emails, phones, mobile devices, they'll still link us together and give us a single digital advertising identifier. So here we understand, to some extent, that the IT corporations work in the interests of law enforcement, constantly stitching our portrait.
So even if a person has changed their phone, they still have the identifier. Through that identifier they can be found in Yandex, Google and other major IT corporations' systems, and surveillance can be set up. What does this give us? It gives us the ability to build a user's portrait: who they are, what they are, how old, what city they live in, what their interests are, marital status and so on. It gives us the ability to track and monitor their movements, to monitor and uncover their social connections and contacts. It gives us the ability to monitor people's presence at a certain time in a certain area, and this is genuinely frightening: if we open the Yandex Audience service and walk along the line of contact, we'll see how wonderfully both Yandex Audience and the marketing platforms of the mobile operators leak information about the number of devices located in a given area, the number of their actual users, broken down by gender, age and interests, and which cities they came from.
And that already exposes actual troop positions. This is used perfectly for guiding precision weapons, and the cases from the start of this year demonstrate how ADINT was used by Israel to guide weapons and attacks on, in particular, a girls' school. Today it's already being used to deliver spyware, thereby widening the attack vector against the device. So this whole marketing business, this whole uncontrolled data collection, on one hand opens up serious opportunities, and on the other hand creates substantial risks. And in conclusion, it's probably worth saying that it makes sense, at least for us in the Russian Federation, to take care of our own security, our own encryption, because if you compare the capabilities for obtaining and collecting data that our domestic services provide, with those provided by, say, some abstract Google, they're incomparable.
The point of entry is incomparable. You can use operators' marketing services, you can use Yandex Audience completely free of charge, after the simplest registration, or not registering in those services at all. Try doing the same in Google, where you have to take part in ad sales auctions, where you have to get special access to mobile app SDK data, where you have to put down a hefty deposit to get that data and work with it. So Yandex and our telecom operators would do well to take note too, kind and wonderful as they are; maybe it's not such a great idea to dump data like this into practically open access.
My colleagues and I have for several years now been building a free browser based on portable Opera that lets you carry out all sorts of investigations. In Telegram, website investigations, cryptocurrency investigations, 1000+ different OSINT sources in there. It's a browser based on Opera Portable, so it stores all its data inside itself. Inside it has neural networks, inside it has social accounts and messengers. All of it runs, if you like, straight off a flash drive, and stores info on all logs and connections there. Perhaps it'll be secure enough, private and useful in your work. So I'll be grateful for any feedback.
Yes, everyone's taken a photo, right. With that, I thank our esteemed organisers. Happy anniversary, success and prosperity. If anyone's interested, my card is on the screen. Thank you.
Q&A
Igor, thank you very much. Your questions, colleagues, by a raised hand. Yep, I see you, coming, coming, coming. You've climbed so far back.
— Igor Sergeyevich, very glad to see you. An excellent talk, as always. Thank you very much. You've been talking about ADINT for over a year now. Sometimes... And yet nothing changes. My questions are about exactly that. The mobile device used to be covered by technical information protection measures. The mobile device mostly passed through those. Now the mobile device is a full-fledged foreign technical intelligence instrument and falls under measures for countering foreign technical intelligence. In your view, is it necessary to give the mobile device a special legal status, so that we can more effectively carry out measures countering foreign technical intelligence?
So you're proposing we go back to the 2000s, when we got a permit from Gossvyaznadzor to buy a mobile phone?
— Some similar arrangement. Well, you understand it yourself. You've just shown in terms of technical intelligence that it works against us, exerts various economic influence, possibly collects information as a whole. Talking about Palantir and the others, everyone knows that perfectly well. About delivering payloads through advertising and targeted ads, malicious payloads, that is, by companies like Rayzone, Sherlock, Paragon, all of that is also already known.
So this is already a very serious threat in the present day. Especially under the conditions of the special military operation. Although mobile phones were banned, they're actively used on that territory. Palantir created the Maven project, a product for tracking. So it's already tipping over in the public information. That percolation point has already been passed. Is it necessary to create this special status for the mobile device?
Well, I'd agree, and that's probably exactly what all the regulators' actions are saying. Regulators all over the world. If I answer your question as a political scientist, I'd say we are living in the age of... There's the Renaissance, and we have the Age of Return. The Age of Return to the old industrial way of life, when states, state entities, held a monopoly on the means of information and informatisation.
So the most correct thing here, of course, is to create our own alternative software and hardware products out of reach of surveillance from outside.
— Colleagues, more questions. Right, hands over there... Ah, I see one, front row.
— Hello, thank you for the talk. Tell me, please, what are us mere mortals to do? Yes, to avoid ADINT, should we stop using Max, VK and Google? No, first of all, install Max. Absolutely. Naturally. First thing. And on your main phone, too. After that it's recommended, well, at least I ran the experiment myself. I tried to strip out as many Google services and automations as I could. But I'm on Android, I've never used Apple, not once, so I'm on Android, and I stripped out all the services, all the automation tools on my everyday phone. Well, there was a period when I had a really cheap, sluggish old phone, and I cleaned everything off it.
It started flying. That's one. The volume of traffic going through it dropped sharply. At the same time, most of the apps, maps and the rest I installed on it simply as downloaded software and downloaded content. Everything worked, it didn't need a constant network connection, especially relevant during the blocking period. Everything was fast, and the phone got quicker. Think about it. It's probably worth doing. If you're still tied to proprietary services from the big IT corporations, it makes sense to gradually start thinking about resetting those ad identifiers, limiting the ability to track you. Naive as it sounds, there's an option in the settings to limit the constant collection about you.
But that's very naive. And you trust those settings? No, of course I don't. Just like I don't trust Yandex when it says we don't collect this, don't analyze it, and we don't have it. When Yandex needs it, it's all there.
— Colleagues, time for one more question. Igor, tell me, you represent a non-commercial security service, there was something about that at the start? The Coordination Council of Non-State Security Structures. The oldest non-profit organization, an advisory body of NGOs in the security field. Doesn't it bother you that you take on law enforcement functions? That's one. Functions that aren't yours. Second. If you're doing, like our colleagues the day before, they said openly: we're a commercial outfit, we collect information about certain individuals, which we then pass on to police officers. So their status is highly questionable, to put it mildly.
And your status here also raises a lot of questions. Have you thought about it? Collecting personal data, there are draconian laws for the legitimate organizations that collect it. And you're collecting it illegally on top.
— Well, it's very debatable that I'm collecting illegally, because it's sitting in the open. Of course, automated processing of even quasi-personal data is questionable here. But on the other hand, answering the question as posed, which, rephrased, is basically "what the hell am I doing here", it goes like this: first of all, I do a great deal of academic work. I have academic papers that were written at the academies of the Investigative Committee, at the Interior Ministry Management Academy, including on topics related to ADINT. So I do practical work here and run research and R&D projects in this field. I work with specialized technical educational institutions, under dual subordination, both to the agency and to the Ministry of Digital Development, where we also produce academic work and teach students and police officers.
So a priori I have to be up to speed on what I'm teaching them. Second point. When reasoned requests come in from them, naturally, I can also apply the practice developed in the course of the academic work, and provide them with that kind of service. Do I do any of this on private commissions? I do. But what I do will always be tied to compliance with the requirements, first and foremost, of the personal data law. We don't violate it. I have almost no doubt about that. Second point. It will be based on the fact that our client has a case file under review, and within the scope of that case file a law enforcement body will approach us.
— Thank you. Colleagues, I see hands, but unfortunately, because of a slight slip in the schedule, there's no time for questions now. So, Igor, thank you very much. Let's send him off with a round of applause. But Igor is with us in the hall today. You'll have a chance to come up and talk with him privately.
Debate “Corporate forensics: arguing about what matters” — Nikita Vyugin (MKO Systems) and guests
Scheduled 11:40–12:10. Guests — Yuri Tikhoglaz, Yan Gorodetsky and Anton Antropov (surnames by ear, see “How to read”).
Moderator's introduction
And now we have a somewhat atypical talk, because, as you've already seen in the program, we're going to talk about corporate forensics. But corporate forensics is a rather ambiguous thing. Sometimes you just can't find a single, unified answer. So today my colleague Nikita Vyugin has brought a whole team with him. And we're going to have a live debate here. Friends, I invite you to the stage. Let's give them a round of applause.
The debate begins: who is on stage
Thanks, Dima.
So, as Dima rightly pointed out, talking on your own about corporate forensics is hard, ambiguous and rather one-sided. So I decided, or rather we decided that we needed to hold a debate. And so, to hold this debate, I invite our guests to the stage. Let me introduce them. This is Yuri Tikhoglaz. First of all, a man with a long forensic background.
Yan Gorodetsky. A man with an equally long forensic past and present. And in order to, let's say, steer, or rather, reason on the topic, of today's debate from the business side, I invite on stage a man with CISO experience, that's probably right to say, and generally with lots of forensics experience from a business standpoint, Anton Antropov.
Right. To kick off the discussion, as the first question, I want to tell you a bit of backstory. Not long ago we ran a training course. Not an ad, by the way, but do come. And one of the trainees was a person from the corporate world. And chatting with me in the smoking area, he said an interesting thing. He says, I did your training. It seems like we need to buy a fleet of those phones that are easy to control, and then we'll be sure that at any moment, if an incident occurs, we'll be able to get the information we need. So, our very first topic is BYOD versus COPE. For those not in the know, I'll explain simply.
BYOD is when we bring our own: I, as an employee, buy myself a phone and use it for work, among other things. It's my personal phone. COPE is the story where my employer provides me with some device, and I use it. Under COPE, all sorts of personal stuff often gets used too, but in the sense that on work devices we can always find some personal things. So, how do I want to structure our debate? I'd like to ask the forensic examiners first to speak for and against one solution or another, and Anton, as a seasoned man from the business side, to have his weighty say on the subject of what practice shows, what works, what doesn't, what's right and what's wrong, from a practical standpoint.
1. BYOD versus COPE
So, Yuri, Yan, who wants to go first? Okay, I'll do it differently. Yan, BYOD or COPE? We're looking at this from the standpoint of usefulness to us as specialists. Yes, from the forensic point of view. From a forensic point of view, of course, corporate devices in that regard, I think, come out ahead. Here we come to the fact that corporate devices are easier to obtain for examination, because we don't have to negotiate with the owner, we don't have to convince him that this is necessary, sometimes for him above all, so we simply go on-site to the client, we get the device through the analysts and analyze it, having all the keys, all the access, without running into the questions of enter the code, unlock it, and if you forgot, try to remember.
Okay, but from a practical standpoint, let me remind you of the COVID times, when COPE practically disappeared, because a huge number of people went remote and there was a massive problem with people hired with personal devices. So from a practical standpoint, in your practice, Yura, and in yours too, how often did you run into BYOD, oops, COPE, I mean? If you take 10 cases, it's still COPE, probably 7 or 8, and 2 or 3 are BYOD, roughly. Well, that's how it is in my practice. Those are absolutely not the numbers that I encounter. Yura?
If we're talking BYOD versus COPE, if it's a properly configured BYOD, then technically it makes no difference, because in the case where a corporate profile is rolled out onto the phone, it's already under the company's management, there's access to it.
If it's the situation where the phone belongs outright to the user, and the company has no means of control over it, then that, in principle, is probably already a security hole. Right.
From an organizational standpoint, yes. It's easier when it's COPE, because the phone went off to the support desk or for replacement, and from there you can work with it. Okay. Anton, anything to add on this topic so far? Well, I'm afraid there won't be much of a debate here, because, forgive the silly joke, "if he hits you, he loves you" is a whim and a fantasy. There is not a single advantage to BYOD that I know of. Unfortunately, our company has this free-for-all going on, but one day I'll get my hands on the CEO's throat and try to fix this situation. Honestly, there's not a single advantage to BYOD. Just not one. You mean from the forensic point of view? From common sense in general.
From the point of view of common sense. Okay, fine. Well, let's even talk money. So what's the advantage of BYOD? You don't have to buy. Buying is the least of it. First, bulk purchasing. Well, the margin there is small, so the benefit is small, but still. We buy in bulk. We maintain a uniform fleet of hardware. That radically lowers the total cost of ownership. I say this as someone who runs into this all the time.
Once the zoo starts, Apple, non-Apple, Linux boxes, and, obviously, standard Windows, you need, we were just discussing this before the panel, you need completely different training for the specialists who deal with all of this. And none of that is free at all. And the cost of hardware, against the annual payroll, is vanishingly small. Unless we're talking about Mac Studios for running the AI on. But those aren't personal machines.
— So am I right that when we talk in terms of COPE, it's not just some custom, well, basically a more or less chosen fleet of devices, roughly identical, so they're easier to control. With some Intune or some MDM system rolled out on them. So that's what we're talking about here, all of that.
Or are there nuances? Well, my colleagues have already covered device control. A corporate device is first and foremost uniformity and centralized management. Control is a part of centralized management. So I'm not adding anything new here. You introduced me as the business guy, so speaking the language of business. First of all, it's simply easier to manage. At the same time, leak control, security functions and so on, they're obviously included in that concept too. And all of it is maintained, roughly speaking, from one place, in the good sense of the word.
— Got it. When it comes to investigations, updating and maintaining the device, on the user's side, that is, in the COPE model, whose responsibility? Is it the user's responsibility to update on time when notified, or does the responsibility rest directly with the company?
It's not the user's responsibility, because it's not their device, that's a), and b), because they don't have any meaningful rights to monitor any non-obvious incidents. If they get beaten up and their phone is taken, that's a different story. Okay.
Anything else to add, maybe? Maybe an interesting case or an example? It's just that, again, I work in the corporate field, and apart from the very largest clients, which are mostly in the central region.
Let's say, the banking sector. Very often the request is that, say, here in the Moscow central branch, everything at the head office is covered by their own devices. You start moving away from the center, and in the regions there are own phones, stories about how the admin mailed a laptop to a remote employee, then took it back, passwords changed, all of that changed, and so on and so forth. My point is that overall, for me, unlike Yan, I have a completely different picture, where out of ten corporate company representatives I meet, only three have COPE. Seven, on the contrary, deal directly with BYOD.
— Well, I can tell an interesting case on this situation. I already shared it with colleagues on the sidelines. So, not long ago we were at a client's, and for context, the client has a turnover of more than a billion rubles a year. A pretty serious company. And we were doing a collection at one of their offices.
As far as I remember, there was a client database leak, and they wanted to know who did it. So we came in, and it turned out to be such a mess it was hard even to look at. I don't even know how to put it, because everyone had their own personal device. A device without any control from the infrastructure, domain policy, they hadn't even heard of that; the computers weren't in a domain. So collecting the devices was complicated by these being the users' personal computers, where they did their business correspondence and kept their documents. And we needed to somehow take them for examination, image them, and then carry on with the case.
And we ran into one of the users simply saying: "Here's the computer, but I forgot the password." Have you run into that? At moments like that you realize that... Oh, and the user was a key one, one of the key people. So, at moments like that you realize that with all its inconveniences COPE is still a must-have. Because if you come to a client who has COPE, everything under policies, everything is great, then, not counting the time spent getting access approved, everything is done almost instantly. But in a situation like the one we got into at this client, well, we tried to solve the issue on our own somehow, that is, we tried to look for workarounds, but in the end nothing worked, because, as they say, there's no trick against BitLocker.
So unfortunately we didn't collect one of the key characters at that point. And potentially, since it was his personal device, he could have simply said "no". Yes, it's his personal computer that he brought to work, because he, well, he had the option to take some basic corporate device, but he said: "I need a better computer. Can I bring my own?" The CEO told him: "Sure, no problem." It's a small office, up to 20 people. And said: "Yes, please, go ahead and use your own computer." The IT department there didn't object at all; it consists of one person. And he calmly did all these corporate operations on his own computer.
Well, of course, he was apparently against it, as we understood. And so he completely cut off all our paths to getting data from that computer.
— Got it. Colleagues, anything to add on the question? Well, in support of the previous speaker. Unfortunately, corporate policies by themselves don't mean they actually work and are configured. It's banal and obvious, but I can't help saying it. I'm a security guy.
And regarding this whole BYOD story. I think it's more a consequence of that whole COVID story, when even in large companies people started working from personal devices, connecting to some VMs or VDS boxes, using personal devices.
Again, as lawyers say, it depends. It depends on what's being examined. What needs to be obtained. And if the scope is only corporate systems, do you even need to seize that device? Do you need access to it? And if you can, for example, get access to corporate correspondence, messengers and so on. There are always alternative ways to get the information you want. It could be backups, if there are any. It could be info from email. It could be information from corporate systems, network shares and so on.
So, the device, sure, the device, but first you need to understand where the information we want to get and need to analyze actually lives. Yes, a micro-comment. They also had email there. Whatever anyone felt like. Mail.ru, Yandex, Rambler. Whatever each person used, they used. Multi-billion turnover, 20 people. Yes, yes, yes. That's a Russian thing. Yeah, we walked around like this for a week afterwards, just facepalming.
And still, to wrap up the BYOD versus COPE topic, again, from my side I see BYOD prevailing so far, again, from my side, that's subjective experience, and nevertheless, giving this system one last chance, or however you'd put it, this policy, overall, would it be right to say that if... If we apply a BYOD policy to our employees, then full compliance with any standards, whatever they may be, is something we can hardly guarantee in reality, in practice.
Or are there still some... You mean technical compliance, forcing a person on their own device to obey strict corporate rules? Yes, yes, yes, yes. Well, obviously there are a number of systems that let you substantially restrict the spread. For example, within the device. So I'd say it's a kind of compromise, because if laptops are bought for people as a means of production, a phone is more of an auxiliary thing, because it doesn't let you generate content well and so on, but it does let you very quickly reach an employee, even when they're on vacation, I know that from experience.
So phones rarely get replaced, let's be honest. I don't run into that often either. Computers are most often corporate. And why am I saying this? Because installing MDM systems, or EMM, whatever, marketing is everything, marketing, sorry. It's a reasonable compromise when containerization is put on the phone, well, not Docker-style, but accordingly, confidential information is stored in some container. That seems a decent approach, at least the mail doesn't wander. Screenshots. Well, screenshots can be taken with a camera from any corporate computer. I'm not really against it, as it were. The systems exist, the market exists. Which tells you that one way or another, the security guys got bent.
So apparently it's not all that scary. But for phones, for computers, for those machines where really large volumes are stored, which is wrong in itself, but you can go far down that road. I don't think that's right. On a computer you can do anything. On a phone too, of course, but less is stored there. The risks are lower. Roughly, it has a right to exist in a corporate environment for phones, but for computers definitely not, only corporate devices with proper control. Right? I don't want to slide into absolutism. Yes, I get it. It's just that from personal experience there was a case. One large organization had a lot of remote employees, including before COVID, because of, what's it called, a teal culture, the computers were often Macs.
And when the sanctions hit, that entire hardware fleet ended up without control. Well, because the organization is large, it's international. And you couldn't violate technical compliance. Nor legal compliance. Several hundred computers were simply left orphaned, because there were no products to control them. Some of them were BYOD. Rolling out DLP on them was practically unrealistic, because forcing a remote employee, well, that's a separate task.
The DLPs available in Russia for Macs that were abroad, the Russian ones, obviously, because of the same compliance issues, they don't work on Macs. I mean, Russian ones, Chinese ones, you have to reboot them four times, disconnect everything, then reconnect it. So, in short, it doesn't work with policies, and even less on BYOD. So it looks more like some kind of shamanic dancing. So the phone is a necessary evil, as I said. There isn't that much info on it, though mail is important of course, but it's in a container, at least. So, absolute protection doesn't exist, we all know that, I don't need to tell you.
But as I said, some kind of reasonable compromise seems to be visible. There's probably room for debate here. Yura, Yan, want to add anything?
2. Should you notify an employee that an investigation has begun
Well, overall, a compromise, we agree. Okay, the second important topic, one I've only run into a couple of times in my practice. Or rather, our users, or people in the hallways, told me about all sorts of interesting nuances. And the topic for discussion: whether or not to notify an employee an investigation has started. So on the one hand this is a philosophical and moral thing, on the other an absolutely practical thing. And, having talked to people from Western companies that left Russia at some point and stopped operating here, some of them had a practice of mandatory notification that an investigation involving them had begun.
I'd like to hear your thoughts on this, for and against. Well, obviously I can churn out the "against" right away. Are there any potential arguments for notifying an employee when an investigation is initiated against them?
— Let me start, I suppose. Go ahead, Yura. First of all, if you do notify, how far in advance? Because, properly speaking, not notifying isn't entirely legal. About two minutes ahead. Right. But you can notify them at the moment they come with their laptop to a meeting with the lawyers, I don't know, the HR director and some management, or you can notify them a day or two in advance, for example, or a month.
— A case from practice. We arrive at a client who received a notice of an internal audit three days earlier.
That notice states that employees are prohibited from destroying documents and information relating to a certain scope. The first thing that greets us is meter-high bags of shredded paper.
— The fence says "don't" too. Yes. Paper that says it must not be destroyed. On the other hand, my favorite interview question for the people I used to hire was this: "You know they're coming for you. You're about to go to a meeting and they'll take your computer, which has hot stuff on it. What will you do?" Many answer: "I'll delete things." What do they look at first? A forensic examiner, getting a device in hand, looks at what was deleted.
So to notify or not to notify, again, depends on the case we're investigating.
Most often, probably, notify, but with the shortest possible lead time, so that maybe a bit of panic kicks in. So that some of those rash actions happen. And then, after that notice, you confront the person with: "Well, dear comrade, right here you violated that notice. Such-and-such data was deleted. Please explain why."
— You're a dangerous man, Yura. Outdated.
— Yan, anything on this? Well, it's hard to argue here. So, on the one hand, a lot of drivers may hate me for this now, but I believe: follow the rules, live honestly, don't break anything, and you won't need headlight flashes about a checkpoint. So in that sense I think you shouldn't notify the person at all, but that's only my opinion, it doesn't apply to the process. I'll say more about that in a moment. For one simple reason. When you try to build the notification process properly, as a rule, the people working the case together with you get into the game. That is, for example, analysts may be working with us forensic examiners, and they have their own view on this, which very often, to some extent, harms a proper investigation.
Because, as Yuri already said, then you come, and your list of deleted files is longer than the list of remaining ones. Yes, it's clear that the anti-forensics issue in any given case may be tiny, but you can't rule it out, you especially can't rule it out under a bring-your-own-device policy, especially if it's all somehow badly, let's not even say badly, just not configured. Now, in the digital age, with our neural networks, it's quite easy to find out how to wipe all the data, all the deleted data, so that afterwards, even by residual magnetization on spinning hard drives, it can't be recovered.
That's not hard these days. So such moments need to be constrained. What time frame is reasonable? Well, one in which the person really can't get started on any deleting. That is, literally: you've come in, sat down with the lawyers, the analysts, notified them, took the device, so they hand it over right in front of you. Because if you give them a chance to delete something, what if they delete it properly. Nobody's saying "dig all you want, we'll find it anyway." No, unfortunately, reality doesn't work that way.
I believe you need to notify within a really, really short time frame, but on the other hand, from a human point of view, the employee must be notified in any case, because it's part of the business process, especially if our forensics isn't spontaneous but some sort of preventive one, then the person definitely has to be notified. And if it's spontaneous, it's stress in any case, for employees, and if the employee has nothing, if they're good, why make them suffer and go through that moral torment.
So I think you do need to notify, but you absolutely have to build in that time window during which the person can't do anything.
— Anton, anything to add? I really do have something to say. But first, a question. Employee tied up?
Good. Seriously, though, what kind of investigation case are we talking about? Because a lot depends on that. I'm not involved in computer forensics, I'm a general-purpose CISO, so to speak. And even that's more in the past by now. But as we know, there's no such thing as a former one. In my practice there were cases with DLP. DLP, if we can count it as a case here, is continuous monitoring. Or it's monitoring that starts on some reasonable suspicion.
And for it to be effective, it's completely obvious warning anyone isn't just unnecessary, it's not allowed. Colleagues said that. But come on, if we're talking Windows: Win+E, Ctrl+A, Shift+Delete, Enter. That's all it takes to destroy a hard drive. Shift. I can write up instructions. For Linux too. Not for Mac. So what am I getting at with all this? To the point that for all this to work, and for there to be no need to choose whether to warn, like, you get it, crow, the question of giving up the cheese doesn't arise.
You need to comply with current legislation. If I remember the trade secrets law correctly, there's a set of internal regulations an employee has to sign when they're hired. Or, if we're rolling out the policies after the fact, we go around and get all those documents re-signed. They spell out perfectly clear things, requirements: how you may and may not use it. And it's stated in no uncertain terms the computer may be monitored. So there's no need to remove anything, just do everything as the law says. Strangely enough, some laws do work. And there are cases, cases successfully taken all the way to court, I mean, well, if you're interested: two young guys were on their way to success, stealing subscriber data and selling on the side, and when it got to court, it took effort, serious effort. It was in the regions, quite a while ago, admittedly.
In short, they were offered a fine or probation. They said, "We've got no money," took probation. Well, everyone who gets the difference realizes that was the wrong choice. Now, as for deleting files, just as a detached comment, it seems to me that in most cases, given the level of trust in the impartiality of the authorities, among the deleted stuff we'll find photos, if not home videos then holiday pictures, and maybe some side gigs, that the person just doesn't want to show, like "I did a little freelance job on this computer." People don't actually steal that often. I'll leave it there.
3. Should you keep a departing employee's device
Okay. Before we move on... Yuri, Yan and Anton all touched on a story that leads into the next question. The only thing that remains a question in this format is notifying or not notifying, but overall it's probably not even about that. So, we've had an incident. We've figured it out, the employee is at fault, we're firing them. Does it make sense, and I'll explain why after, to keep their device, not wipe it and hand it to the next employee, but store that device for some time, because last year a man came to me and said he had cases from a corporate financial environment going back 8 years. That is, he's got that time depth, he's got a number of devices where the investigation was supposedly finished, but then something happened, and it turned out to be connected, and he had to bruteforce BitLocker on machines eight years old.
I don't know how it ended, it's very interesting, but in general: you've finished the investigation, you wipe it and hand it to the next hire, or does it still make sense to bury it somewhere for a while, or back it up?
If anyone has thoughts on this, go ahead, Yuri. Depends on the case, because I know, for example, that at the job before last my laptop is still in a safe, not wiped. Simply because there are situations where deleting data like that can, among other things, mean a fine from the regulator that's significant for the company.
If we understand we're never going to court over this particular employee, and the information is preserved in a forensic image that's properly documented, then maybe we don't need to keep it.
Well, from my side I'll add: in any case, at least in my work, we always create an image of the device. So keeping the device itself, just holding it in a safe and not handing it on to the next employee, I don't see much point, we won't extract anything more from it. You switch it off, the RAM is gone, and we already have the image, so that's it, nothing left in it. So the image needs to be kept for a certain amount of time, law enforcement has it written down, we also try to keep it all six months. And again, as Yuri said, it depends on the case. Some cases, where clients ask us to keep the images longer, on the assumption that in the future this may be connected to some other matters as well.
The image is stored by practice, that is, best practice is onto a hard drive and into a safe, and let it sit as long as needed. Anton, anything to add? Well, honestly, I don't see why you'd keep the laptop, not the data from it. Maybe there are cases I'm not aware of. As for 8 years: the statute for especially serious economic crimes is 10 years, so it makes sense. Okay, got it, thanks.
4. Proactive forensics versus reactive
And now, what we've already started touching on in passing. Proactive forensics versus reactive forensics. What to bet on? Why "what to bet on"? Because, unfortunately, in the real world, very often in companies that aren't big IT or financial giants, people have to choose. Either we bet on monitoring everything all the time, surrounding ourselves with tools, and, well, by then there's nothing important left there anyway. Broken?
— Some technical glitches, never mind. So the first option: we monitor all the time and are constantly in a state of readiness, while keeping a stock of tools, and a staff of people who service all this, watch it, collect, analyse. Or the second option: an incident happens, and either on our own or with hired teams, well-known ones, some of whom I know are here, we call them and say, "Mate, I'm in trouble, come and help, I'll pay."
One doesn't rule out the other, we all get that basically, but overall, the case for reactive and for proactive forensics. When is one better applied, and when the other? Bearing in mind that we don't live in an ideal world and the budget isn't unlimited.
Anton, you're smiling very slyly, go on, I can see it. Well, I just like to talk, I think everyone's figured that out by now. It's as if proactive forensics doesn't exist, because forensics is investigation. It can't be proactive, the term itself is an oxymoron. But, again, from personal practice: covering the whole organisation with DLP is unrealistic, because even for large banks these are substantial costs that don't pay off.
So yes, obviously, you need to narrow it to focus groups and so on, and that can probably count as some kind of proactive practice. In general, I think that in the real world it makes sense to act incrementally, as always, step by step.
— First cover a small number of people, then, if necessary, if you can clearly see problems are happening, expand the coverage, so you're not just burning money but delivering real benefit, obviously.
There was one more thought, I've forgotten it, I'll come back to it. Yan, from your point of view? Well, I think that, as Anton said, proactive forensics is really something from infosec, because it's closer to prevention. Whereas reactive forensics is already, you might say, mitigation. So if we go back to the whole point of the forensics procedure, I lean towards reactive here, because, first of all, it's more interesting. Because if a person is constantly in prevention mode, they either start screwing up less or hide the traces better. That's not interesting.
Like, you look into the computer, and there's less there than there could be. So, in my practice, there simply is no proactive forensics, because when the next case comes along, we grab the equipment and run off to extract data. That is, we run specifically to look at what happened, where it came from, why it happened at all, who's to blame. As for prevention, I think that's more something the infosec folks deal with, or some separate, dedicated department of in-house forensic examiners, or maybe someone uses outsourcing for this. So it's more, as I said, a matter of information security inside the company.
Overall, yes, the question isn't quite right, I probably didn't phrase it correctly. What I meant was the story of preliminary and ongoing collection of information for the purpose of prevention and, let's say, essentially, prophylactic, to understand where an incident might potentially start. Not when it's already happened, and you've been called in and you run to extract. Exactly, that's closer to infosec, indeed. Yes, here the tasks of proactive examiners would be very similar to how a DLP system works. Yuri?
Proactive forensics, from my point of view, is something strange, like fortune-telling. You can't conduct an investigation in advance without conducting one. Here it's probably more a question of readiness to conduct investigations. That is, understanding where the data is that you'll have to work with if something happens. Backups of the information on computers, corporate systems. And an understanding among employees, IT people, security guys, of what to do when an investigation is needed.
— Right now this seems to sit on the side of having regulations ready and monitoring users' compliance with those regulations. Documenting systems. Yes, understanding where the information you'll need is located, who's responsible for that information, how to get access to it. Then, when the need arises for reactive forensics, everything goes somewhat faster and easier. That's really what it's about. Reactive forensics goes easier, than... If the client has done a certain amount of preparatory work so these investigations can be conducted.
I'll be honest, from my own subjective experience, this now sounds like Alexander Dmitriev's take when we were sitting at that meetup. And he says, "Look, in general, it's all written down, there are golden rules." And I asked him in reply, I said, "Sasha, how many companies in your whole life and in your work as a pentester have you seen that had these written, where the rules were written by the Golden Book?" He answered: less than 1%. And my point is, about what you're saying, we should know, and we should have procedures for responding to an incident. That's clear. How often is expectation, well, expectation vs reality, how often is the data actually where we expect it to be, especially if it does turn out to be an incident?
Twice in my practice.
— Well, there's your answer. You're a lucky man. A lucky man indeed. On that, I've remembered a thought. On proactive defense, or rather proactive forensics, whatever you call it. The intimidation approach works relatively well. It's a very controversial approach in itself, but an internal company-wide mailing saying, guys, from Monday we're monitoring all of you, significantly reduces potential incidents, at least for a while. Unbelievable, but true. Everyone lies low. Yes, that's a well-known thing.
Well then, do you have anything else to add? Just a small aside about how best to live so that even reactive forensics gives greater results. You started discussing this with Yuri. You don't have to look far, we have this wonderful procedure called e-discovery. There are the first two or three points there, it's enough to follow them, and any reactive forensics will go like clockwork, if those points are carried out properly and precisely. That is, we will clearly define the information, identify it and define the scope of where it's used and located. And everything will be fine.
Good reactive forensics still starts with preparation.
— Absolutely. Including regulatory.
A question from the floor and the close of the debate
Dima, tell me please, do we have a Q&A section in this? If anyone has questions, we'll gladly let our guests answer. Anything to add, Igor Evgenievich?
— I have a question for your guests along these lines. What is the field of application of their knowledge, their services? Well, obviously, small organizations simply can't bring you in to solve their problems and will solve them themselves. And big ones, like Sber, they can just buy specialists themselves, that's exactly what they do. A level of qualification sufficient to avoid bringing in outsiders.
If I may, where is the area in which you feel you are in demand? If we broaden forensics to include external attacks, then personally I talk three or four times a year to companies, let's say, if not small, then mid-size, I'm not ready to throw out figures, I won't go into detail. These are small infrastructures, a few dozen servers that got encrypted, for example. They come and tearfully beg for help recovering and figuring out the causes.
So far, in 100% of cases I've talked them out of it. Because there's no point in trying to restore a 100% encrypted infrastructure. In most cases there's no point trying to buy the keys. Especially now, when it's not encryption but wiping, for geopolitical reasons. So the question could, I'd say, be broadened. Who needs this and what to do. But here I'll hand over to more specialized colleagues.
In my practice forensics, digital forensics specifically, goes hand in hand with the discovery procedure, for one simple reason. Because the electronic document disclosure procedure isn't as developed here as in the West, but still. Why? Because you need to know how to properly present electronic evidence to the regulator, to the court, at the request of some law enforcement agency, and so on. So in a corporate investigation we use the algorithms of the discovery procedure in order to ensure something like, probably, the points of Federal Law 73 on forensic expert activity. That's completeness, integrity and, probably, reproducibility, because if we present evidence unsubstantiated, it won't be accepted as evidence.
Fulfilling these points lets us give our analysts a full-fledged report they can take anywhere. So we currently exist at the junction of these two fields, in this way, even given the current situation.
And still, to answer your question, who needs this and why: what, for example, the Big Four companies do, is supporting clients in, say, major international arbitrations. There the clients are, say, lawyers in major international arbitrations, or, for example, corporate investigations, also in large companies, where the result is also submitted to court in one form or another. Because to fire an employee, conducting a forensic investigation doesn't always make sense. It's probably easier to run it through some HR mechanisms. And sometimes it's easier to just fire them. Yes, like that.
— Thank you very much. Nikita, we're out of time. Got it. Thank you very much. Yuri, Yan, Anton, thank you so much, it was interesting. If there are any questions, we're still here for now. Thanks. Thank you all.
Colleagues, I have a small announcement. Now we'll take a break and come back to the hall at 1:10 p.m. Meanwhile you can grab a bite, have a smoke. I also remind you about the booth area. And now we'll put up a QR code. That's the information portal Valeria Mikhailovna told you about yesterday. Here you can find articles by authors who, basically, took part in our conference. There are a full seven of them. And also, if you're a client of MKO Systems, you can get extended access to the portal, where there'll be our articles, training videos and also test cases that you can use both for training and simply for investigation.
Have a look, play around. Please.
Vladislav Azersky (F6) — “DFIR vs WSL: digital forensics where two worlds meet”
Scheduled 12:15–12:35.
Moderator's introduction
One p.m., as I said, we're opening the second part of today's event.
[A break is cut from the recording (22.5 minutes). By this point the program had shifted: Vladislav Azersky's talk was scheduled for 12:15, and the moderator opens the second part of the day with the words “It is one o’clock…”.]
And Vladislav Azersky will open it. For a long time, Windows and Linux existed for the specialist as two worlds. But with the arrival of WSL, the boundary between them has become much more blurred. What all this changes for the DFIR specialist and how it works is exactly what our colleague will tell us. Let's give him a round of applause.
Vladislav, the floor is yours.
Talk
— Right, I hope you can hear me. Yes. Today we'll mainly talk about a component that appeared at a certain point in time, called Windows Subsystem for Linux. And I honestly wanted to talk about some kind of forensics here. Right, next. Ah, there, all good.
So let's talk about it. But I'll say right away that there won't be any bone-crushing digital forensics here, because, as it turned out, there are no forensic artifacts specifically related to this component. Very surprisingly, there are no ETW providers that would push events into the Windows Event Logs. And that's surprising. You start googling, not even googling, you download the extracted manifests, you search them for occurrences of Subsystem, Linux, and it's all empty. That's the first problem. Second: okay, are there any text logs, or anything else that might be generated.
Maybe it'll be sitting in some SQLite database, or somewhere else. No, empty. Of course, there are still certain traces left we can use to determine it. There are some mentions, again, in those same event logs, but they are more indirect. And in the registry.
As for the history of WSL's development, one thing to note right away is that there are two versions. The first appeared in 2016, the second in 2019. On the next slide I'll say more specifically how they differ. But for the end user the underlying principle doesn't really matter much. So, what do we actually need WSL for? If you still prefer working in a Linux environment, but at work you've got a Windows host, or you're a developer who mostly works on the Windows operating system, then this might be a solution for you, because you can easily set up a certain component in Windows, flip a few switches, so to speak, and then download a distribution, open it and work from the console without any problems, a proper Linux console.
As for the next turns, so to speak, not of development but of history, EasyWSL is already helping us, a kind of tool that's open, open source, free. It lets us take images that we can pull from Docker Hub, and convert them into a full-fledged WSL distribution. So now we at least get the ability to import something, and also export it. And in 2025 this happened: Microsoft decided to finally release the code of the kernel and everything else publicly. But there are a few points: some drivers and libraries there, again, are still proprietary, but the bulk of the code is present.
Now let's... Let's look at what WSL looks like and the fundamental difference. Maybe for a forensic examiner this won't be terribly important, and the same goes for some developer and everyone else, but it's good to know these subtleties so you understand later how, for example, an attacker or a user can interact with it.
WSL1 and WSL2: how they work and what changes for DFIR
In the first version there was simply a layer, a translation layer, which essentially implemented Linux system calls passing through the Windows kernel. And the second version, that's the one we'll discuss. We can see a big picture here on the right. And basically, in principle, this is what it all looks like. A Linux virtual machine is spun up, and inside it we have each distribution, because we can create several of them. For example, we want Kali Linux and Ubuntu on our Windows system. Consequently, we'll have two of these dotted boxes inside that virtual machine. And so we can already interact with them somehow.
In the end we get this: there's a host box running Windows. At startup, the WSL Service starts inside it, the Linux VM. And inside that, when we want to open some distribution and work in it, naturally, a WSL distribution gets created. And as for the specifics of working with WSL, there's a lot of important stuff, because in some attacks and cases, of which, honestly, there were quite few, but mostly, if you look at what's in the news field, you could find something. About the specifics. Let's start. First. We can run Windows executables from inside WSL. For example, we simply go into some distribution, get a bash shell and can run things, either by specifying the full path, or by just typing, say, notepad.exe, and it will launch.
Great. And here it's worth noting, for example for a SOC specialist, that in this case the parent process will be wsl.exe, which in turn launches notepad.exe. Next, there's the fact that every distribution has the Windows file system mounted, that is, the host's. And we can easily, as a two-way thing, both copy and save certain files. We'll take a quick look at that a bit later too. Again, for analysts there's an important point, if, again, they have telemetry, that all the file operations we perform from WSL, for example, we downloaded some file and want to drop it into the Windows file system, will be handled by the system process dllhost.
Why do I emphasize this? Sometimes it happens that in some EDR solutions there are simply no detection rules for this behavior. And for the most part this second paragraph applies more to WSL 2. The first one has a similar mechanism too, but it works a bit differently. And the last one is network isolation. Essentially, WSL 2 is fully isolated, that is, its network stack. And in this case what can that lead to? If our Windows host system has something like Sysmon installed, which collects telemetry on network connections, or an EDR, then later we won't see anything, because all of it will be happening inside the virtual machine that hosts our Linux distribution.
That's why attackers in some cases took advantage of exactly this. Next, let's walk through what all this actually looks like. For example, we can run commands in a distribution, specifying it from Windows OS. Here we just use the native tool wsl.exe. For example, we do a listing to see how many distributions we have. And then we specify a particular one of them. Next, the root user and exec. As you've noticed, there's no situation here where a password is prompted, or we even pass one here. This means that in the course of working with the WSL tool, it turns out that into every Linux distribution installed on the system we can easily log in as root.
That's a particular quirk of how it works. Consequently, even if we've locked down root somewhere, so nobody could do anything, from Windows we can easily pull it off anyway. And the second point I mentioned, the peculiarity, is that once we're inside a distribution, in WSL, we can launch, say, calc.exe by specifying its absolute path, and it pops up right here. All of this works by default. This exact mechanism, where some Windows executable gets launched from a Linux distribution. Here we can see how to disable it. We simply write into each distribution, it has a config file, /etc/wsl.conf.
For example, these two lines. Then nothing like that will happen anymore. Second point. As you noticed, in the second column we had the fact that we can copy files back and forth. And what does that look like here? If we want to copy from Windows into some WSL distribution, we go to the UNC path \wsl$ and then specify the name of the distribution itself. And then we get the files and directories that are valid for the Linux OS. And the second point. We're inside the Linux distro and we want to copy some files to the host box. Here everything is already mounted. It lives in the /mnt directory.
We go into, say, a folder and can copy right away. Here at least a simple listing has been done. Okay, so why did I talk about these peculiarities? Because these peculiarities are already being used in some cases that you can see publicly. But before that I want to say that even when you're gathering some information to understand which groups did what, when they did it, and whether they even did it at all, you still need to double-check the info, because several sources mentioned that a number of groups, such as Turla, FIN7, Ryuk, used WSL. This was in two or three sources, and it was very strange, because there were no other mentions anywhere.
It turned out to be just some kind of neural slop, where the person didn't actually check what was written, so this here is more up-to-date info.
Cases: Qilin and npm typosquatting
The first mention was actually around 2017, but more or less one of the companies wrote up that we found an ELF file that interacted with WSL. There are no specifics, none at all. They just found some particular sample, analyzed it, said we found it, and everything else, goodbye. The second one, essentially, is some affiliates of the Qilin group. Early February 2026, the attackers gain access to the infrastructure in the standard way through some RDP, or there's already some RAT there. Next. I don't know why, but they decided to pull this off. They either activate or deploy WSL. But for this we actually need to already have full control over the machine.
If we're talking about the situation where WSL isn't there. What do they do next? They load the ransomware in ELF format and run it in the WSL environment. And what happens? Since the Windows file system is mounted for us, all the user files will be affected, because this ELF file will go through the files already located on NTFS. And what do we end up with? Encrypted files on the host Windows OS. And the last recent story. There was a certain campaign that the attackers pulled off. It was in August. What's the specifics? It's npm typosquatting. They created around 40, or even 50, npm packages with similar names. That is, the difference could be that they swapped two adjacent characters, or, for example, replaced an l with a one.
And after such a package was delivered to the system, a script was run there. What did this script do? It first checked whether the system, that is, whether the user is running under Linux. If they're running under Linux, then it checked the next thing. Whether in this case they are running specifically within a WSL distribution. For this it checked two environment variables, because they're always there by default. And what happened next? The ransomware was downloaded, copied from WSL to the Windows file system, and launched. Oops, not even ransomware, but a stealer. And it stole credential data. These are basically the cases that you can at least find.
Attack vector. Let's move on, we narrow it a bit, that is, we even generalize. What do we end up with? We have two options for the attack vector. The first is when the attackers gained access to the WSL distribution via a malicious script or package, like the third of the news items presented that were published. And the Windows side, when the attacker gained access to the host machine. This can be in a corporate infrastructure. A person decided to connect, raised their privileges and started working. And here two scenarios matter. First, WSL isn't installed for us, therefore you'll need to get full control over the system. That means we need admin rights.
If they're not there, we won't install WSL, because you need to enable exactly these two components, one related to WSL and one to the Hyper-V system, that is, to virtualization. If it's installed, we can then download some distribution and work with it further.
Persistence, delivering a distro, and detection
Let's look at persistence. We have, again, in each distribution a file /etc/wsl.conf, there's a boot section, there's a variable, say, the word command, a key. Here we can specify a command that will be run at the launch of the distribution itself. Here we actually just have a reverse shell, which gives the attacker the victim's shell. And here we'll go with the scenario that will most likely be more applicable. That WSL is installed on the system, they connected to the host and then carried out this whole thing. So they did a listing, then specified a certain distribution and then did the next thing. They wrote a command in the /etc/wsl.conf file.
The second point, very specific, which mostly won't work, is persistence via a WSL plugin. There is such a possibility, but here we need a digital signature. And most likely issued by some trusted authority, because otherwise for a test one we'll have to disable a certain mechanism where Windows doesn't trust self-signed digital signatures. Great. Let's say attackers can sometimes steal them from some company, and they signed it. Next, what do they need to do? To sign the DLL itself that has some payload, for example, download some beacon and run it, and then register it at a certain registry path.
And here we just specify the path where all this will be. And here there are two ways it works. We can restart WSL, or shut down the computer. Essentially, at system startup, or when the service restarts, the payload located in the DLL will be launched. And one more very important point. Let's talk about the fact that these distros have to be sourced from somewhere. We can, for example, deliver them.
There's a project on GitHub that, essentially, lets you create a custom WSL distribution and add a payload to it. In this case a calculator was used. The person, the developer of this script, wrote the code, and basically here's how it looks. We specify there the name our distribution will have, and -p calc.exe is which executable file will be launched at startup, exactly, when you enter the distribution itself.
Next, this calc.wsl file is created for us. We deliver it to the host, specify this command. It calmly shows up here in the listing. And we launch it. On launch, wsl -d calc, the calculator launches in WSL for us. It's built, most likely, on a fairly lightweight Alpine Linux image. And there it will be somewhere around 10 megabytes. How does this look if we need to check for their presence? First, in the registry, at a certain path, it may be hard to see here, we have information, how many distributions we have installed. Here there's a path, here there's a name, here there's, and also additionally, if needed, what the name of the virtual disk of this Linux distro that appears for us is.
And then we go by certain indirect logs that we can somehow dig out. There are two journals, I won't talk about them in detail, because what matters to us here is exactly the entries we see here. First, where our virtual disk is located. Most often, by default, when creating a new distribution it'll be called ext4.vhdx. Consequently, we can somehow think it through and in one case set up rules with a SOC, in the other search.
Even simpler. We can, say, skip those odd things the attacker does, or some user who wanted to deliver it this way in the previous case. We, for example, put together some small 10-megabyte distro and made an archive out of it. And what can we do? We can simply deliver that archive, import it and specify the path where our ext4.vhdx virtual disk will live, and then it starts up just fine. Another example. Okay, we don't want to deliver anything from our own server. We can use the winget utility, which will grab something from the Microsoft Store, for example, a Kali Linux image again, or from the public open winget repository.
Here, for example, the attacker runs winget search kali, finds the distros, and then depending on the choice it will be downloaded either from the winget repository or from the Microsoft Store. And here we have to keep in mind that if the attacker uses the winget utility, then in this directory, under this file pattern, we'll find a lot of useful stuff. The command that was run will be recorded. So in those cases we'll see that winget search kali is present, whether it succeeded or failed. And the last one, that an installation was performed.
I'd also like to mention, though I won't go into it in much depth, that the SpecterOps folks did some pretty good research where they showed they can launch processes directly in the Linux container without using the WSL utility, as we saw earlier. A specific COM interface is used for that. If you're interested, you can read it, it's all laid out in a lot of detail there. They got to this after the WSL source code was published.
OK, so we've gone through what WSL is and what its specifics are, but what about digital forensics, really? Well, as we can see, not much gets generated, there are really only some indirect links, so we come back to the standard. Since our host system is Windows, this is Windows forensics. And then, if we're working with distros, we look at what we can find, what forensic artifacts, or logs there are in the Linux distro itself. And here we can build a sort of pipeline of actions. First, we go through the base artifacts and logs of Windows, which let us establish whether WSL is installed, whether it's in use, that is, whether the components are on and there are already some distros.
That's without context, just to find that out. Then we move on to the digital forensics of the Linux OS. If we noticed we have some distros, that they appeared at some odd moment, that they aren't used by a developer or some user, then what do we do? We either do a triage, or copy that ext4.vhdx image and then analyze it. Within DFIR, I would proceed specifically as an incident responder, from the position of forming a hypothesis based on current attacker TTPs, because most often we're dealing with some external threat, which often operates by its own particular standard. And then I would switch to analyzing the forensic artifacts and logs of the Linux OS.
If you're interested in reading some other research, and generally getting this presentation, head to this Telegram channel, everything will be there, most likely either today or tomorrow. Thank you. I'll take your questions now.
Q&A
— Raise your hand, colleagues. Any questions?
Yes, I see you, I see you, coming, coming, coming. That's great, because I thought there'd be no questions again.
— Vladislav, good afternoon. Very good talk. And my colleague and I are actually thinking about a question right now: what if, for example, you set up auditing on this Linux subsystem and output the logs through a path mounted on the Windows system? Is that an option for monitoring activity in there? Yes, it's possible, and most likely that's one of the options, because I didn't mention it here, but during my research I found there are these plugin systems, we can somehow use them, and I noticed: oh, there's a Windows Defender plugin, I can install it and something happens. I'm looking at it thinking, OK, I hope it's not purely tied to their MDE, and it'll at least save some logs to files.
Turned out it doesn't. But potentially you could indeed look at WSL plugins, but there's a big problem here. You'll most likely have to disable Secure Boot and grant Windows trust so that a self-signed certificate works. In part, this functionality they introduced could work very well, because it would run inside that Linux VM, and you could simply collect right away from the various distros that get spun up in there, some kind of logs.
Otherwise, yes, most likely in this story you either just restrict users from using WSL, monitor the enabling of the components, because, again, only a limited number of people will use it, hardly some accountant or lawyer. But if you collect from there, you'll have to, inside the distro itself, set up some tool for monitoring and data collection, so that later, as a good option, you copy it to the host machine and pull it from there. Well, it seems that's really the only viable option so far. Thank you.
— Colleagues.
— Thanks for the talk. Here's my question. Don't modern EDRs look inside Hyper-V containers? WSL mostly runs on Hyper-V, right? Yes. Honestly, I don't have that kind of analytics or statistics. But I've seen somewhere that at least various international vendors are introducing that functionality. As for the Russian market now, I don't know, but maybe some have already at least done something basic to look inside, or are on their way. Because after all, if you pay attention, it's 2026, and the main cases appeared right here. Apparently there is some need after all to dig deeper into this Linux VM.
And essentially those same plugins, again, if they could be installed properly, could at least be some kind of solution.
— Well, the mere fact of Hyper-V starting up is effectively a signal too, because... It will be a signal, 100%, but again, you have to take into account the context, and, again, I'm not in favor of using WSL there, it will be used by a small number of people, so most likely you just target those users who genuinely need this tool. Can you encrypt that virtual disk using WSL itself? I mean, roughly speaking, Kali runs on an already encrypted disk, and when the forensic examiner comes, accordingly, he can't do anything without some key, and the key, accordingly, is fetched over DNS or something like that.
Potentially you can do that, but attackers are unlikely to, because most often they need, without any potential errors, to do something further, because what comes next? You come into the chat to talk with them, and they go: OK, great, send us some files to check. And then it turns out that if they can't be recovered, or something else, that's it, that's already a big problem. They won't get their potential money.
— All right, okay, thanks. Colleagues, any more questions? I don't see a single hand. Vladislav, thank you very much. As always, thorough.
Yuri Pavensky (independent expert) — “Examining electronic storage media when information security tools are unavailable”
Scheduled 13:10–13:40.
Moderator's introduction
We're moving on. And, as you all know, in an ideal situation experts always have all sorts of software tools to extract everything, to work with everything, do it all right and neatly. In short, as it should be. But, as we know, situations are far from always ideal. What to do in that case is what our next speaker, Yuri Pavensky, will tell us. Please welcome him with a round of applause.
— So, what about the presentation?
— Vladislav got carried away and walked off with the clicker, so I run around. Here you go. Thank you.
Talk
So, good afternoon, dear colleagues, dear participants of Moscow Forensics Day 2026. First of all, I'd like to thank the organisers for the invitation and the chance to speak at today's tenth anniversary conference. My name is Yuri Alekseevich Pavensky, I'm speaking here as an independent expert. And today I suggest we work out together how to examine electronic storage media in a situation where the corporate data network is no longer functioning properly, and the usual access to security tools and their telemetry is no longer available.
I'm sure many of you are wondering how such a situation could even arise in practice. After all, we're such a great company, right, we have a huge number of different security tools at our disposal, including certified ones, well, for example, things like SIEM, EDR, DLP, TI, Web Application Firewall, NGFW, antivirus software and many, many others. On top of that, we regularly run pentests, hands-on exercises, we improve our incident response processes and we go through a huge number of audits by various regulators. And again, here you'll probably say, but the results of those audits always get the highest possible marks, so what could threaten us?
It would seem that with that many tools and measures any incident will be detected in time, and the information needed to investigate it will always be right at hand for the information security staff. But, unfortunately, this assumption contains one of the fundamental mistakes. Because no matter how modern and strong your information security system is in practice, it in no way guarantees that, if you face a serious cyberattack, you will necessarily keep access to your usual security tools and their telemetry. And again, you're probably asking yourself, so what do you do in a situation like that, when the corporate network is down, and the usual access to the security tool consoles is gone.
Let's look into this question using a specific incident as an example.
A case: from a leaked database to encryption and extortion
Let's say a database of a well-known online store was published on the public internet, which included customers' personal data, email addresses and the passwords needed to log in to their accounts. That database included an employee who worked at the company developing a well-known IDM system, who used his personal email address as his login.
I think you all understand perfectly well that as soon as any database is published on the public internet, attackers, as a rule, always work with that kind of information. Basically, that's what happened here too. The attackers, using the discovered password, later managed to get access to this employee's personal email and public cloud service. What happened next? Next, using open sources of information, the attackers obtained some contact details about this employee, identified additional personal email addresses, as well as his corporate email address, and also established the employee's position and organisation.
They didn't stop there, and later, while analysing the contents of the compromised resources, that is, his personal email and cloud service, they established info about the customer organisations this employee regularly worked with, information about privileged AD accounts, and also information about the hosts used and the remote connection methods.
Next, the attackers continued the email correspondence with one of the customers and under the pretext of carrying out additional work obtained remote access to one of the organisation's jump hosts. What happened next? During the first week the attackers behaved as if they really were a contractor, didn't raise any suspicions, but after some time they slowly started gathering information about the domain infrastructure, about the specifics of the network architecture, about accounts and rights, and also about critical servers and databases.
All of this subsequently led to the compromise of several AD accounts, to obtaining the information needed for further use of various privileged and service accounts, and in the end they secured a persistent presence in the IT infrastructure of that organisation. But, again, they didn't stop there. The databases they discovered, containing information about finances, employees, customers and logistics operations, they exfiltrated.
And at the same time, while analysing the organisation's workstations, they came upon the workstation of an IT department employee, which held an unencrypted backup of a personal Apple-made mobile device. Naturally, they liked it a lot, they took it, and later, from parsing that backup, they pulled out photos of confidential documents, pulled out the email correspondence stored there, including some that contained information about credentials that had been passed to one of the contractors.
And also, accordingly, there were network diagrams in there, and some other accounts that were kept in the notes. After the data collection stage was over, the attackers moved on to destructive actions, which ended with the destruction of the IT servers' backups, the deletion of the discovered config files of core-level switches, well, not only core-level switches, but basically other switching and routing equipment as well.
And after that, corresponding changes were made to the config of the core-level switches. And all this led to very sad consequences for the organisation. The corporate network stopped functioning properly, its connectivity broke down, and this led to consequences such as degradation of the domain environment, loss of normal access to the security tools, and certain centralised restrictions on the corresponding telemetry.
Yes, and most importantly, on the previous slide, I forgot to mention, that already at this stage, after these consequences had set in, a rapid reconstruction of the attack with standard tools was no longer possible. Next, accordingly, the attackers, well, not even next, I'd say, at the same time as these actions, the attackers changed the access passwords to various databases, after which those databases were encrypted.
Some time later the attackers made contact with the IT department employee via the Telegram messenger, informing him of the compromise of the organisation's IT infrastructure, as well as of a ransom demand for restoring access to the data. In addition, the attackers said they had captured his unencrypted backup, which contained sensitive data, and they also threatened to publish the full package of information online. As you understand, such a package would clearly have damaged the honour, dignity and even business reputation of this employee. And to confirm the reality of that threat, a certain fragment of the data was in fact published in one of the public Telegram channels.
So, let's draw a brief conclusion from this case. As you may have noticed, that the compromise of a single personal account can lead to an incident of this scale. Furthermore, giving contractor organisations remote access without proper control can lead to quite high risks. Beyond that, you also need to bear in mind that unauthorised storage of confidential information on personal mobile devices, as well as in their backups, can cause damage both to the company itself, and to the organisation's employees.
And probably the saddest and most dismal thing is that in conditions where we don't have access to the security tools we're all used to, when we're simply, in principle, not accustomed to such situations in practice and don't even expect them, the only key instrument for restoring the full picture of the incident is exclusively digital forensics.
So, well, thus, yes, having discussed, accordingly, all this information, well, we all understand perfectly, yes, that at this stage the entire attack chain has been established. Well, only we can see that whole attack chain. But in practice, at the moment such an incident is detected, the information security staff don't see that picture yet. All they have is the initial data, such as the corporate network not working normally, failures of specific nodes and databases.
It may be a message received from the attacker, and, accordingly, it may be a fragment of data published on the public internet. That's exactly why, within the response process for this kind of incident, there are ten interrelated stages, from recording the initial report and preliminarily establishing the scale of the incident, through to reconstructing the attack sequence, restoring the IT infrastructure and closing the incident. And, again, I'd like to note that we won't be going through each stage in the most detail today, due to the limited time of my talk, so, if you're interested, yes, and you really want to understand what needs to be done in this situation in the most detail, yes, at each stage, you can scan this little QR code, so, in the first folder there'll be a document, it's called "General Response Procedure", yes.
You can download it, this QR code will also be on all the following slides, so at any convenient moment you'll be able to get access to that information. But if someone doesn't manage to scan it, you can come up to me separately, and I'll give you that information. So, well, again, from here on I'll focus on the core logic of choosing actions toward each node in an incident like this, I'll talk about how to properly preserve volatile artifacts and obtain verifiable data copies. Along with that, we'll also carry out the reconstruction of the attack sequence together.
Isolate, image or leave alone: the logic for each node
So, speaking directly about choosing what to do toward each node, here you need to bear the following in mind. That here there's no such thing as a so-called universal rule. Something like, we absolutely must, with you, immediately isolate every host in our organisation. Or, for example, that we are strictly required, from every host, without exception, to collect absolutely all the information in order to have some kind of complete picture.
That is, I believe these rules overall should be flexible and not categorical. Why? Well, because if we, with you, isolate every host, well, you know, if we, and here I'm still considering a more or less normal case, if some system of ours happens to still be alive after the incident, then after isolation there's a risk that something might then go wrong with it. Especially with some database that, say, data was being copied into at that moment. Or, for example, if it's some, you know, insanely complex incident to investigate, we risk simply and thoughtlessly losing the necessary artifacts that would have let us reconstruct the whole attack chain.
If we're talking about full data collection from every host, let's imagine our organisation has 2,000 hosts, 10,000 hosts, 100,000 hosts. Can you imagine how long we'll be collecting a full data package from every device? I'm afraid we'll be collecting information until our company shuts down because of this incident. So that's also a completely wrong approach, you definitely mustn't do that.
So, you'll probably ask, what are we to do, then, given these nuances? It's all very simple. We need to assess the node's connection to the incident. That's the simplest thing you need to do. If we've established such a connection, the next step is to assess whether the damage is ongoing, both on the host itself and through it, or whether there's any suspicious network connection. If any one of these conditions is present, then we move on. We check whether any network access remains to any systems and resources that are still alive. Again, yes, you'll ask, but our corporate network is down, what access to other resources could there be?
Well, actually, that's very simple too. If our box lives in the same VLAN as those systems, there's a good chance that such access to those resources may well remain. So you also need to keep that in mind. And the last general rule, which applies to everything. It's also important to account for technical limitations, because it may turn out that the box, say, isn't running quite stably. And if you collect data from it in a not-quite-correct way, there's a risk the box will simply stop working. Or, for example, take the same switches and routers. Take Huawei, for instance, let's keep it simple, just as an example.
As you well understand, there's a huge number of models by purpose. The devices can differ from each other. And it might seem there could be one and the same operating system, Huawei VRP, but in practice different commands may actually be used. So they, in turn, when executed, may again have a different impact on those same network devices. So all of that also needs to be kept in mind. So, moving on to specifics. Once again, if we have a connection between the node and the incident, but at the same time no damage is ongoing, there are no suspicious network connections of any kind, if at the same time no network path remains to our resources, then in that case you don't need to drop everything and isolate that host.
Well, you can of course do it if you're very worried, very scared, I don't know, you can cross yourself if you like, but you definitely shouldn't, well, at least I wouldn't do it, because the host, essentially, is already isolated, it has no access to anything at all, in any way, it's essentially running standalone. So in that case we just need to preserve the volatile data, and then not touch that host at all, until we get to obtaining persistent data. Not at all. If we have a host that's connected to the incident, but at the same time active encryption, modification, deletion, copying of data is going on, caused by the attacker, well, probably the first thing you'll say is, that's it, now this host definitely has to be isolated.
No, colleagues, here too this question, I think, is rather philosophical, yes, and making that decision, you have to manage it, because, let me explain, because if, for example, it's some critically important host, yes, if, for example, some operation is currently running against those same databases, there's a risk that we do network isolation, the connection drops, some commands don't complete, who knows, something somewhere, we don't know what scripts the attacker might be using at that moment. And the database itself could go down.
And restoring it, especially in conditions where the AD servers' backups have been deleted, will be problematic. That's an example. So, that's why, if we understand that we have at least two minutes, colleagues, more probably isn't needed. If we understand that in those two minutes no catastrophe will happen to that node, under those conditions we can obtain the volatile data.
Because if, again, we do network isolation, we'll lose a huge number of such artifacts. That's 100%. So while they're there, we have to make use of them. Again, I'm not telling you that, you know, using Linux as an example, I'm not saying you now need to enter about 67 or more commands to obtain all the necessary volatile data. No, of course, you should by no means do that manually.
You won't fit it into two minutes then. For that, it's recommended to use an appropriate automated script. If you don't have one, consider that after today's talk you do. You can also scan the little QR code. There, first of all, you'll find a full list of commands that need to be entered to obtain some top-priority volatile data. Depending on the operating system: the Windows family, macOS, Unix-like operating systems based on the Linux kernel.
I'll get to switches and routers a bit later, but that information is basically there too. And the corresponding automated scripts are there, which you can also use in practice. They're PowerShell ones, you can... Well, some are PowerShell, some are other kinds of scripts. In short, you can use them in practice. In a couple of minutes all the necessary information will be exported for you and a corresponding report will even be generated.
Right, what else did I want to say. If our node is in no way connected to the incident, that node must not be touched under any circumstances. Why would we need it, why waste extra time on it? Let it keep living its own life over there. Now, if during the incident investigation we do somehow end up at that node, then we'll need it, then we'll actually work with it. But as long as it's not connected to the incident, we don't need it. If the node is powered off, there's no need to power it on and boot the preinstalled operating system for this.
Well, at least not within this stage, definitely not... In fact, in principle, you shouldn't do that at all, because then our persistent data will change. So, and the last thing I'll probably mention is virtual machines. So, if we have the technical ability, and if we're confident that the hypervisor cannot make any unacceptable changes to the guest operating system's data, in that case we can create a snapshot preserving the RAM. And I want to say up front that before pressing that little button, creating the snapshot, please don't do anything on that box whatsoever.
So that everything is preserved intact and safe. Again, you'll surely say, what hypervisor could there be here, if our corporate network is down, how would we even connect to it. But actually, that's very simple too. If the servers aren't outsourced, somewhere on the other side of Russia, in another country, but are in your office, of course you can take a non-domain workstation in the form of a laptop and connect either to the corresponding SAN switch, or directly to the management port of that equipment. Then in that case you'll get access to the hypervisor and then, accordingly, to the virtual machine itself.
But, accordingly, the other specifics, which have to do with encrypted drives, with switches and routers directly, in terms of choosing the logic of action, I mean, in terms of some unfinished remote sessions and so on, you can also study all that information in detail by following the link in that QR code.
The order in which volatile data is collected
So, we've now reached the stage itself, which is called the order of collecting volatile artifacts and obtaining data at high risk of loss. So, what needs to be kept in mind here? That at this particular stage, first and foremost, we're not interested in all the info that exists on the media as a whole. I'm talking about this exact stage, we're primarily interested in the data that may somehow be altered or lost as a result of the OS running, as a result of the attacker's actions, or as a result of network isolation. What are we interested in first? First of all, we're interested in basic node information, like the hostname, details of the current user account, the version of the operating system or the loaded kernel, the system date, time, time zone, the time from an independent source.
If there are any time discrepancies, those discrepancies also need to be recorded. And we're also interested in the node's uptime since the operating system booted. Next, we're interested in information about user activity. Here we can talk about such information as data on active users and sessions.
We're interested in cached authentication data for the current sessions, information about running processes and services with their launch parameters, and information about open files. Next, we're interested in local logs with a high risk of being overwritten. You'll probably say, why collect local logs at this stage, they're stored in persistent storage, so we'll take a sector-by-sector, or, as some people say, bit-by-bit copy, and we'll somehow recover those logs from there, somehow pull them out. But you know, colleagues, if something is written there constantly, every second, then again, over that time there's, in general, a risk, especially while you're also still busy collecting some volatile data, that some of the data may get overwritten anyway.
So, if possible, some local logs also need to be exported. You don't need to export all of it. It happens that local logs weigh in at terabytes. And then you'll be exporting until night, I don't know, for a week, probably, exporting, I don't know how long. But, again, we're interested in the info for at least the past 24 hours. I don't think that will be some really huge volume of data.
So, again, I'm talking about logs, and you'll probably ask which logs I mean. Well, if we're talking about Windows family operating systems, for my part I would recommend still exporting part of the data from logs such as System, Security, WinRM Operational, TS-RCM Operational, TS-LSM Operational, OpenSSH Operational, I don't remember if I said it, WinRM Operational. And, most importantly of course, Windows PowerShell and PowerShell Operational. If we're talking about Linux operating systems, in that case we're interested in logs such as syslog, audit, secure, kern, messages, daemon, xrdp and xrdp-sesman. Sesman. Oh, and auth as well, one of the most essential logs.
After that, ah, if we're talking about macOS, here we'll primarily be interested in syslog and Unified. So, what next, yes, what do we do after the logs? Once we've exported those logs, we're interested in information about loaded drivers, kernel extension modules and, accordingly, information about active inter-process communication mechanisms. And the very last thing that should also interest us is information about the host's network state. That is, information about its network interfaces, active established network connections, listening ports, the routing table and information about neighboring hosts. Here I mean ARP and NDP.
As I said, we need all this data. Again, under no circumstances do we save it to the persistent drive, we save it either to external media, or to an isolated network storage, that operates independently of the main corporate network. Well, again, that's if someone happens to have such an option, who knows.
Switches and routers: how to connect and what to collect
So, on this slide we'll discuss collecting volatile data from switches and routers. So, what else needs to be said here? In general, in principle, a loss of connectivity in the corporate network, of course, by no means yet indicates that our network device has somehow been compromised. No, by no means does it indicate that yet. It only indicates that either a technical failure occurred on the equipment, which happens quite often.
It could be the clumsy hands of the IT department staff, who at some point were setting up the corresponding configuration. Or it could indeed be attackers. As it was in this particular case. So, how do we connect to such network devices? Clearly, if there's no connectivity in the corporate network, if we can't somehow connect to these devices over the network, the only correct option, probably, is to connect from a standalone, non-domain laptop using the appropriate console cable.
And then the very first thing you must do before you establish the session is to start recording that very terminal session, so that all further actions are captured, and to establish the time from an independent source. Next, the very first command after you enter your credentials should be the command to view the shell history for all users. Why is that so important? It's so important because later, when you use other commands to obtain volatile data, there's a high risk that old entries will be pushed out by new ones.
So it's better to capture that whole history first. And then, accordingly, what information do we need to obtain. It's all very similar here, actually. Information, that is, the hostname, information about the current account we're working under.
That's the system date, time, time zone. And again its deviation from an independent source. Then it could be the operating system's uptime. The operating system version, it could be local logs, administrative sessions, network routes and whatever data remains in the main operational tables.
What else did I want to tell you about switches. Ah, yes, most importantly, of course, when collecting volatile data we never, under any circumstances, use write commands. That is, again, what I want to say is we use not just commands for viewing information, but commands verified in advance, whose execution should be obvious to us in terms of what result we'll get and what load on the device it may cause. Because, again, speaking of the load on the device, there are nuances there too. That is, again, you have to understand that any connection to a network device, I don't know, any command entered, even a read-only one, has a negative effect on the shell history, on the corresponding local logs, and also, accordingly, it may also be reflected in the AAA systems.
And most importantly, it can cause a short-term increase in load on the device itself. Again, speaking of collecting information, you'll probably say: well, here we go, again you have to enter a huge number of commands manually, again, that'll take a lot of time, and why is it all needed? Agreed, you don't need to enter all these commands manually. That is, again, for such actions you use an appropriate automated script, which I've personally used myself. In maybe one or two minutes it will export all the necessary information for you.
And by the way, it's also available at this link, via the QR code.
Persistent data and verifiable copies
So, now we've moved on to persistent data. I don't know why the letter O is glowing red. Oh well. So, what I want to say here. That is, within this stage, our task is to obtain such a copy of the data whose origin and integrity could subsequently be verified. Yes, for that we record the data source, the method of extracting the information, the tool and its version, we note the storage location, well, where the obtained result is saved, the time of the operation, and we calculate the corresponding checksums.
So, again, yes, what...
[Here the recording breaks off mid-sentence: only part of Yuri Pavensky’s talk made it into the stream, and there is no Q&A for it in the recording. The rest of day 2 (Maxim Sukhanov, Alexander Dmitriev, Viktor Alyushin, Artur Igityan and the closing of the conference) is not in the recording.]