Moderator's introduction
And we're back for the second part of today's event. While the guests gather in the hall, I'll say a bit about what our next talk is about. It used to be that the main task was to get as much data as possible from various digital sources. But when you have several terabytes of various photos, videos and countless other data that today's devices store, absolutely any of them, including mobile ones now, and the big question becomes: how do you find something important, and work with all of it? And here it's going to be quite hard to get by without automation. How to move from creating forensic images to AI analysis of multimedia will be presented by our colleagues and conference partners from ELETEK, Vladimir Greshnov and Sergey Savinkov.
Guys, the floor is yours. Welcome them with applause.
Introduction: what ELETEK does
Good afternoon. We'll probably wait another minute for everyone to sit down.
— Right, now a short introduction. Our company does both integration projects and development of our own software and hardware tools. Our main areas, so to speak: we make specialized systems, hardware-software complexes, a secure execution system, software for collecting and analyzing data, tools for forensic data acquisition, operational acquisition and so on, and, accordingly, some large server-based, integration systems built to the customer's order.
Today our talk will be devoted to two of the topics listed. These are tools for forensic data acquisition and tools for analyzing heterogeneous data. For the first part I'll hand over to my colleague Vladimir, and I'll come back.
Vladimir Greshnov: the duplicator and the compact copier
— Yes, well then, greetings to everyone. As Sergey already said, today our company is represented at this event by two of its divisions. I'm going to tell you about the hardware and the software that we need for acquiring data and for working with data images.
The first system, we spent quite a long time developing it, I'd say. Last year we already presented it for the first time. But this year, or rather this month, it's going into series production and is starting to ship. This system of ours is a duplicator. A duplicator that lets us quickly and, most importantly, safely acquire data from the drives under examination.
We need it both in lab conditions, that is, you can work with drives that have been seized, and in the field, say, on some kind of site visit. This system runs off, it's powered from 220-volt mains. It can also be powered from a power bank if needed. So that's an option too, so that, say, in the field you can easily power it up and work with it. Its main purpose is copying specifically from external data storage devices, such as HDDs, SSDs and flash drives.
It's also possible to copy NVMe drives through an NVMe-to-USB adapter. Inside there's a full-fledged microcomputer. Lately we've started working very actively with modular solutions. That is, we take a certain type of brains, roughly speaking, from a microcomputer. Then we ourselves lay out boards with the ports we need, built-in hardware write blockers and so on. And we connect them together, and we end up with a full-fledged device.
As for copying. It can copy either pass-through, that is, from the drives under examination to some trusted drive of yours, or to internal storage. The internal storage is a fast NVMe disk. In terms of capacity, right now we mostly ship two-terabyte disks, but at the customer's request we can add 4-terabyte disks as well. There's no problem there, thankfully all of that is produced and actively sold now.
Our next product is our most compact copier. Basically, many of you have already heard quite a bit about it, some may have worked with it, as this product is supplied to the Interior Ministry, that's the main customer, for about the second year now.
This copier lets you copy data specifically from flash drives. Why is it limited to flash drives here? Originally it was developed as a field-operations solution, that is, officers from operational units asked for the most compact, self-contained solution you could take with you, one that would literally fit in your fist, fit in your pocket. As I said, it's self-contained, meaning it's powered by an internal 2000 mAh battery.
That's about 2 hours of copying, if you're acquiring from drives that support the 3.0 interface. If the drives are 2.0, then accordingly it will run longer.
Inside it we also have a full-fledged built-in microcomputer that makes it work. As storage here we use a microSD card, so after copying you can simply remove that card. Unfortunately, the second box isn't shown in the photo here. The box of this product, which is where that card sits. The card is removed, connected through an adapter to a computer, and there you go, and you can work with the copied data.
As for copying modes, it supports sector-by-sector copying, file copying, copying by masks, but mainly it was intended for copying files specifically. Because copying files is faster for us in some kind of operational conditions. We kept sector copying too, since a duplicator should still be able to make a sector-by-sector copy. On the next slide you can see the app for controlling both the first and the second system, though even here I'll qualify that: not so much control as pre-configuration.
On the second screenshot you can see that the operating mode is selected for the system first, that is, which mode it will copy in. Sector-by-sector, file-based, or by masks. The first screenshot is our first screen, the mode is monitored, that is, the copying itself is tracked: what's happening, what status it's in, how much space is left on internal storage. And the last, third screenshot, here we work directly with the dumps themselves or the file containers that were copied. They can either be exported. Exported to the mobile device the app is running on. Or, accordingly, you can work with that data directly on the drives they were copied to.
SATA and USB 3.0 write blockers, software systems
The next product we make, it's actually a whole group of products. These are our write blockers. Well, basically, devices everyone is very familiar with, the devices themselves are pretty simple. So here we have a SATA write blocker. We've been producing it for... quite a few years now.
Design-wise it's as simple as it gets. It can work with 2.5-inch HDDs, and 3.5-inch ones when you hook up additional power. With SSDs, basically with all devices that support the SATA interface.
And the second write blocker is our new solution. It's a USB 3.0 blocker. So it hasn't... hasn't gone on sale yet. That is, right now it's just finishing, so to speak, its final tests. Basically, it's also quite a simple device. You connect power, it powers up. It connects to the computer through port two. And through the USB Type-A port you hook up any flash device you like. If needed you can also hook up an NVMe drive through an adapter. But here, unfortunately... for now we don't support maximum NVMe speeds on this device. After all, we intended it more for flash drives.
The next group of products we make, these are no longer hardware-software systems, these are our software solutions proper. The first solution is a system we call Element-P. It lets us acquire data live from computers with Windows file systems, Windows and Linux. So how does it work? This program sits on a drive or on a flash stick. So, we connect it to the computer. The program launches, you create a task in it for the copy types you need. Say you want to grab absolutely all the files from the file system. Go ahead, you create a task, run it, and all the files just fly over to your device.
But what's interesting about it is that you don't have to grab all the files, you can search for specific files by extension, files by path, or files by MIME type. We started working with MIME types in our duplicators not that long ago. So it can detect if an attacker has stripped the extension off the file you need, or given it a different one. So the program will understand that these documents... are, say, not an audio file but a text file. And the program will grab it.
This slide shows screenshots of the interface. So, roughly what it looks like and how it works.
And our next software system is a program for working directly with data images. So with the previous solutions you can make a sector-by-sector image. And with this program you can take that image and... load it in. See what files are inside. Apply various sorting to those files by extension, by modification date, creation date and so on. And this program also supports recovering deleted information. That is, you can run a scan for any deleted files or deleted partitions.
The program can also show you that such-and-such partition used such-and-such file system, and in the deleted one, this other file system. And, accordingly, try to recover as much of the deleted data as possible. It can also work with signatures. That is, it can do signature-based search. And search by MIME types, accordingly. Well, on this slide you can see the actual screenshots of the program.
Now, this program and the previous program for copying, they're already at the final testing stage right now. So basically, in the near future you'll be able to get demo versions to try them out, accordingly, on your own setup. Well then, the block I've been talking about is finished. I'll hand over to my colleague.
Sergey Savinkov: analysing mixed data — text, audio, video
Next up is a talk on the analysis of heterogeneous data.
I'll tell you about our suite of software products, which we also integrate into systems built to the customer's spec. It has a kind of modular architecture. That is, there's a set of processors and a set of data collectors for gathering data from internet or internal portals. That could be some closed internal systems deployed on your side. Data can be combined both from the internet and from internal sources, land in a nominally closed perimeter, be analysed there, and provide the tooling.
Tooling for multi-user work by analysts. That is, these are multi-user portals with various access controls, where the analysts themselves or other staff search for information, organise it, and put together reports. Well, today's talk will mostly cover the central block, the data processing. If you're interested, we can tell you about the rest at the stand. So, the various kinds of data, roughly speaking, we split into text, audio and visual data.
We've been working with text for about 20 years. Currently we use classical algorithms for precise classification. That is, when you know in advance what you want to look for. These can be fairly abstract topics. For example, texts relating to physics, or to extremism. Anything at all. Or there can be more specialised classifiers. That is, find me specific factual mentions in the text, names, company names, addresses, assess the tone, the type of text, meaning scientific literature, free-form statements, media, and so on.
As a rule, classifiers are applied as a set. That is, specialised pipelines are built. A specific sequence of classifiers that gives the analytical result you need. That is, find me such-and-such a topic, but with a mention of, say, a particular region or country, and with a negative tone, for example, or a popular-science tone. AI tools are also now widely used for text analysis. They already let you search after the fact, that is, you first index your large sets of text data extracted from various sources. Let me repeat, this can be a single archive of data extracted from a phone, from flash drives, from computers, plus data added from open sources, internet sites.
And then you search them per a task, for example, identifying all the passages of text that mention such-and-such semantics with the required tone or behaviour, behavioural type. Audio data, what gets added here? Essentially, audio transcription gets added, a task which on the one hand is already solved these days, but on the other always needs an individual approach. That is, audio data recorded with noise, in foreign languages, with many speakers talking, with constant switching between languages, requires extra tuning and fine-tuning. We have a solution for that. You can fine-tune on your data, you can feed it your own subject dictionary.
That is, when there's specialised vocabulary, off-the-shelf transcription tools cope worse. They try to pull the transcription towards a more common word. Whereas you might have some jargon or specialised vocabulary. Accordingly, the next step is identifying speakers. For now we only do speaker separation. That is, speaker 1, speaker 2. Speaker identification is something some of our partners do. We haven't gone into that area yet. Then, accordingly, the transcribed text can be translated from foreign languages and a summary produced. That is, if you have a long recording or a set of recordings, you can get a single summary of what was discussed.
Or ask specifically, for example, whether there were any facts in this conversation or not. Visual data. On top of the first two types, an image is added. Essentially, the input is video or images. If it's video, it's split into frames and a set of tasks is run. Essentially, this is working with the frame as a whole. There's classification of what the frame shows: vehicles, people, landscapes, military, documents, and so on.
Faces are handled separately. That is, searching for specified faces, or simply clustering all the faces present in the video or images. Specialised search for specified emblems, flags, signs, certain objects, and so on. That is, you define the semantics and the frames where those semantics are present get sorted and picked out. Accordingly, the output is a large array of indexed video data, with markers placed where the objects you're looking for appear: faces, again texts, images.
And separately, we've done fairly deep work on recognising various documents. That is, these can be both printed and handwritten documents. Naturally, the quality is a bit lower for handwritten. But printed documents can also come in very different formats. That is, I'll show examples now. So, more detail on audio recordings. The specifics of our solutions. We can process recordings of any length. Streamed recordings.
There's stable memory consumption. That is, we don't try to pull the whole recording into memory at once and so on, like some solutions do. Multimodality. This means you can have a mixed recording of dialogue among several people who speak different languages at the same time. The processor picks out individual segments and, based on the language, routes them to the appropriate recogniser, the text transcriber.
Adaptability. As I mentioned, you can specify your own specialised dictionary, your domain. Which greatly improves recognition quality when you have terminology or slang. That is, you don't have to pre-train on those recordings, you just slip it a little dictionary of terms. For example, you have some radio intercepts of surveyors, with lots of terminology. And then the transcription will go very... the quality improves several times over just because of that. And now, accordingly, some examples of local analysis. That is, solutions range from a laptop up to a server cluster. For example, we have a laptop solution where processing is done in a single language.
That is, at any given moment you're processing in one language. If your recording is multilingual, processing will just take a little longer. Because it switches between languages. On a workstation you can already support multiple languages in parallel. And the speed goes up. And then, depending on the data volume, there's the server solution...
Here are visualisation examples from the portal that stores the processed archive. Well, here it's specifically about video data. But basically, the indexing and visualisation solutions are designed for all data types at once. So once you've processed all this multi-format data, you then search across all of it as in a single portal. That is, you enter queries and it searches video, audio and text data alike.
Video processing. Let me dwell on it once more. It's searching for the frames you need, classifying them, searching for faces, known or unknown. That is, it can simply process your archive, a large archive of videos, and lay out clusters of people, showing that such-and-such people appeared in it. And there they are in different, well, in different clothes, at different ages. Simply where they were, whether they were together in the same video. Beyond that, the range of analytical tasks is limited only by your imagination.
Special timestamps are created so you can jump straight to the face, emblem or statement you need. There's an editor for preparing report materials. That is, you can cut out the segments you need and make a final report video. And accompany it with a text report on another page. And another important thing is processing video not for transcription, when people are talking in it, but to describe what's happening in the video. Especially if the recordings are long or nobody is talking. Here too the range of tasks is unlimited. It's like getting a brief summary of, say, a nine-hour video where nothing happens.
From the first hour to the eighth you've got trees, wind. A fence, and five cars drive past. And from the eighth hour people come in, something's happening. So this often cuts down the analyst's working time. When they read a summary like that, they either see they need to jump to that section, or see right away there's nothing of interest. It works well, for example, on predefined topics. So when we're looking for videos on hacker topics, videos of suicides or fights. Or we're simply searching, again, when you have some street camera surveillance footage, and in 24 hours two cars drove by, at such a time, such a time, such a time, and at such a time a group of people walked past.
So, again, out of 24 hours we get three two-minute fragments that need to be analyzed in more depth.
Image classification: these can be pictures as well as, strictly speaking, video frames. Again, the groups are defined based on your tasks. They can be generic ones like portrait, landscape, maps, or some specialized ones, for example, markings on a photo, or detecting documents, printed or handwritten documents, official documents, or just documents lying on a desk, and so on.
I mentioned document recognition as a separate block because the task is so specific. We've also worked on large sets of documents, both printed and handwritten. We got pretty good results and quality. Again, it's designed for cases where, say, you... Well, no need to even mention large archives. Even the data pulled from a phone is already a huge data set, which I hope nobody's trying to go through by eye anymore, because 200 gigabytes of photos and files is just impossible to look through. Unless you have a lot of free time.
The next block is another example of the same specific task with photos, that is, for example, describing what's happening in a photo, a screenshot, an avatar, and so on.
Working with text data, analysis of large volumes of text data. Here it's shown on open data from the internet. Again, the same thing: data pulled from a phone, a flash drive, a computer. These can be full-text documents, but also chat-type documents. There's classification, author extraction, link analysis. If factual data is mentioned somewhere, that is names, organization names, addresses, phone numbers and so on. And a reporting portal that lets you save your queries on top of that and add things manually.
A huge archive of indexed texts, tens of terabytes. It lets you find the information you need in seconds, and narrow down by certain attributes. Time, as an example here. File type, where it was found, the extraction it was found in, and so on. Examples, as I mentioned, of classical text classification algorithms, which we still use today in parallel with AI. Because on some tasks, when you have a rigidly defined category, they work both faster and with better quality.
In fact, as a rule, we build classifiers for specific tasks, for the customer. We already have a large body of ready-made classifiers for standard topics. Even specialized ones we've already implemented. We have dedicated linguists for that, refining them and so on. As a rule, like I said, they work as a set.
An example of visualizing heterogeneous data. Again, a portal where data of all the types I've listed flows together. And you search through it, and it no longer matters where it came from: a closed source, some disk, or the internet. The archive lets you search everything at once. And it's multi-user. Each analyst can have their own set of classifiers. So someone works on topic 1, someone on topic 2. When they log in under their account, they see the material picked for their topic. On top of that they can refine it, add another query, narrow it by time, pick out materials for the quarterly report, the current report, the annual report and so on.
So it's a system of collections. Fairly standard modes for analytics portals by now. Accordingly, presenting the information as cards, text, maps, graphs and so on. Jumping, as in the video data example, into a specialized editor for the data type. All of that is implemented. Again, as a rule, when we're automating, say, some department, there's fine-tuning for that department's needs.
We talk about how the work is organized now, how people do it now, by hand or on some other system. And basically all these visual forms get tailored to the customer.
Q&A
That's it in brief. If there are questions, we're ready to answer.
— Colleagues, thank you very much. Your questions. Yes, I see one.
— Hello, thanks for the talk. On the first block, I'd like to clarify a point. There's a need to reproduce normal operating conditions on devices. Is there a write blocker that can fully pass through the device, the source, so it can be connected to the personal computer or laptop itself and powered on, when what's needed from the blocker is not passing through the controller's data but the data of the source drive itself, meaning the interface, serial number, SMART and so on. Because ordinary blockers just pass through the data of the controller the blocker maker builds in, the blocker's own.
Here that particular thing is partially implemented and is being implemented further in the new USB 3.0 blocker. So in the SATA blocker we don't support that. And the second question, about MIME types. Will it be possible with your product to detect, for example, archives or crypto containers disguised as some data types, images, video, some common formats, like GGUF files posing as a language model, disguised, well, like crypto containers named with a GGUF extension. Will that be possible? Yes, that will be possible. Here's how we've implemented MIME types in general. We have certain groups of MIME types that we initially set up for users.
Say, audio, archives, as you already said, images. And we also have a custom group available, where you can add all the file MIME types you need. And I'd also note here that sometimes you need the ability to search for all encrypted types possible on the device. Would it be possible to implement a search for encrypted archives disguised as some common file formats? Honestly, we haven't worked in that direction yet. So that's something to discuss with the developers specifically. But I hear you, we'll definitely raise the topic. On the second block, I'd like to clarify how this is handled.
With a local model on a rig deployed at the examiners' site? Or is this some kind of online analysis? As a rule, all these systems are deployed locally. I said we build a comprehensive solution that can even be two-tier. An open segment and a closed segment. And all of that runs directly on your side. That's why I said data from open and internal sources can be combined in one archive. Naturally, not in the cloud. This is a question about specs. What specs does the rig need to have to run the model in working conditions? And how is the context issue handled? I mean, we understand that the main problem right now is the context size, which tops out at 1 billion tokens.
How did you solve that problem? And what models are used? With what parameters? And how do you make sure that after a large data set the model doesn't start getting confused in the system prompt and in the internal instructions and so on. The hardware is chosen by the task, naturally. That is, how much you want to process and how fast.
I gave an example from a laptop on up. As for all the other questions, how should I put it. Again, the model searches the whole archive. So if you have a terabyte, it's already indexed. It has a schema of that archive. And it searches within the context depth, that is, whatever you ask it. But it's also limited by the standard things. So as a rule, it's, for example, two GPUs that run a model of 38 or 72 in size. So there's nothing special going on there. So the query context is, naturally, limited. But for analyzing an archive that's already structured, with a schema built for it, that's quite enough. So it performs the search across the whole archive.
Colleagues, any more questions? Let's do it by raised hand.
Could you tell me, please, about the USB blocker. What's the technical solution? Is it a microcomputer in there too, or a signal processor? Look, we use a microcomputer there. And on top of that, our modular solution. We don't use the computer itself, just the brain of it. Then our own board is laid out, which connects to it. And that's how it works. So the blocking is still done in software? Inside the device, yes. Thank you.
— I see a hand, coming over. Can you keep it up? Because I lose it among the rows. Thank you.
— Hello. A question about the Element VD software product. Can you name its key advantages compared to competitors? Or is it simply an alternative because it's Russian software? Well, first, yes, what you said. It's specifically Russian software. And probably the main advantage it has is that it runs on Astra Linux. And working with, well, I'm not sure now, R-Studio. It seems to work with E01.
That was our main, large-scale task. Working with E01 files on computers running the Astra Linux operating system. Got it, thank you.
— Thank you. Right, colleagues, we have time for one more question. Slava, I see you. Alright, fine, two. Okay, colleague, you go first. You raised first. Yes, thanks for the talk. I wanted to ask, when you came to visit us, I already asked you then how you handle interruptions while making an E01 image. Back then you said you weren't aware of such solutions. I pointed out that they're already on the market, and you wanted to look into adding them. So if we're creating an E01 image and the drive drops out, everything stops, with you it all had to be started over, but there are products that can resume.
Well, if you tell us which product that is, we'll at least be able to see roughly how it works, because, based on the information we have, the E01 format itself simply doesn't support interruptions, since checksums are calculated within the format itself, and after an interruption that can't be done right. I already told you: Belkasoft, Tableau, you can look them up. Does Tableau really have that? Yes.
We'll look into it, alright.
— Right, good afternoon. Tell me, please, how is the emotional state assessed? You had that written on your slide. Well, the emotional state by the tone. So you mean audio recordings? No, no, you had it for audio. For audio recordings. Well, by timbre, naturally. By timbre, pauses. Well, some people just speak louder, some quieter, and so on. No, but it does assess it. Was there scientific work, or just a guess?
Scientific work was done, there are already standard solutions for this. Well, I see, okay. And how will it be described, in what words, a fight, say? Waving arms, or something else, because someone standing there doing something, is that enough? Well, it assesses the dynamics of the video, so if, roughly speaking, it describes that a fight is taking place, because... Well, no, it could be hitting, or something else, you see? No, well, it doesn't break down the type of fight inside, it just says a group of people... That's not what I mean. What's it for? Say, a large number of video images that may contain some unlawful actions.
And if it outputs that for you as text, then you can search by that text too. Yes, yes. But what do you search for? Well... Exactly. What words does it write it in, so it's clear... Look, as a rule, as a rule, I had an example there, a query isn't written as a single word, but as a group of words with, well, special linguistic terms. And when an analysis portal like this is set up, a query library is built.
Because the words written can vary. It may just be people moving actively, approaching each other, right, so in fact the word "strike" may not appear explicitly. "Fight," because it may not be there, they may be standing close to each other, shoving, and someone just stabbed somebody with a knife, for example. And that may be written out explicitly. But on top there's the linguistic query, which must account for all these nuances, like, well, for example, worked out on a test sample, right. Well, so first you need to study what you output, then think about what... Well, not what we output, but what we want to search for.
And then the linguists themselves tune the queries for that. So it won't be just one word or two, it'll be something like a mini-classifier, in effect. No, I understand, it's just that usually linguists don't do this, but computer experts. No, it's better with linguists, because they better understand the semantics of the operators, how to link the words. And one more thing, in what form is the dictionary built, for specific terms? Is it just text? Or do you still have to pronounce the word plus text? No, no, no, it's text.
In these kinds of audio recordings you can run into specific words like that, right? I mean, certain terms, slang, right? Because with standard transcription it may just not have that specialized word in its context. It's more likely to recognize it as some more common word. When you specify that there are nicknames, or, I don't know, special terms, it will prioritize looking for those, yes. Yes, you just enter a word. A word, Enter, a word, a phrase.
— The closing one, sorry, that's ours.
— Testing, testing. A follow-up to the previous question. I'd like to know how you've implemented reproducibility of results. A common problem with AI-based analysis is that the output results change. How have you solved that? Obviously, you can lower the temperature, the model's tendency to make things up. How have you implemented reproducibility? So that, for example, a video is first analyzed one way and you get certain results, and then sometimes the model starts giving different results after analysis. And how have you implemented that? This is very important for expert examination and analysis results. Processing the same video twice, or different videos?
Yes, yes, yes. The parameter there is hard-set. So the reproducibility will be very high. Overall. It's like this making things up, the imagination of the neural net. With this kind of analysis it's practically... Well, it's not switched off, you can't switch it off, right? But it's heavily restricted. Because otherwise, when transcribing audio or video, for example, you'd get arbitrary text resembling something. So we restrict the parameters strictly there. Well, we do preprocessing, slicing. I think, for audio tracks, right? And accordingly, its flights of fancy are very heavily limited, yes. So each language has its own set of models. For example, for slicing and preprocessing a video file there are separate small models.
So there's a whole stack of models.
— Colleagues, thank you very much. The guys have a booth here. And you can talk to them all day today. Let's send them off with applause. You can leave it there, please. Thank you.