Інтелектуальна система підтримки прийняття рішень на основі нейронних мереж для виявлення медіаманіпуляцій

dc.contributor.advisorКасьянов, Павло Олегович
dc.contributor.authorБаздирев, Антон Андрійович
dc.date.accessioned2026-07-10T08:19:43Z
dc.date.available2026-07-10T08:19:43Z
dc.date.issued2026
dc.description.abstractБаздирев А. Інтелектуальна система підтримки прийняття рішень на основі нейронних мереж для виявлення медіаманіпуляцій. – Кваліфікаційна наукова праця на правах рукопису. Дисертація на здобуття наукового ступеня доктора філософії за спеціальністю 124 Системний аналіз (12 – Інформаційні технології). – Національний технічний університет України «Київський політехнічний інститут імені Ігоря Сікорського», Київ, 2026. У сучасному інформаційному середовищі медіаманіпуляції є одним із ключових інструментів інформаційного впливу, що особливо загострилося в умовах російсько-української війни, масового поширення повідомлень у соціальних мережах, месенджерах і цифрових медіа, а також високої швидкості тиражування маніпулятивних наративів у різних текстових варіаціях. Для українського інформаційного простору ця проблема є не лише науковою, а й безпосередньо прикладною, оскільки своєчасне виявлення маніпулятивного впливу має значення для безпекової аналітики, OSINT-досліджень, моніторингу інформаційних операцій, фактчекінгу та медіаграмотності. Разом з тим, у практичній роботі аналітика недостатньо лише визначити, чи є певне повідомлення маніпулятивним. Потрібно також зрозуміти, яку саме техніку використано, які фрагменти повідомлення є тригерними, які зовнішні свідчення підтверджують або спростовують висновок системи, як безпечно сформулювати нейтральний виклад і які випадки слід розглядати в першу чергу за умов обмеженого ресурсу уваги. У зв’язку з цим у дисертаційній роботі розглядається задача побудови інтелектуальної системи підтримки прийняття рішень для виявлення медіаманіпуляцій. Запропонована постановка поєднує методи системного аналізу, нейромережевого опрацювання природної мови, інформаційного пошуку, безпечної генерації тексту та прикладного реранкінгу випадків. Центральною ідеєю роботи є перехід від окремих моделей до цілісної доказово-орієнтованої архітектури, у якій результати нейромережевих модулів інтерпретуються як елементи аналітичного контуру, що завершується практично придатною підтримкою рішень. У першому розділі запропоновано трирівневу постановку задачі. Перший рівень відповідає за аналіз окремого повідомлення і включає визначення ризику маніпулятивного впливу, багатоміткову класифікацію технік та локалізацію фрагментів-індикаторів. Другий рівень реалізує доказово-орієнтований контур, у межах якого здійснюється добування свідчень, повторне впорядкування знайдених кандидатів, типізація доказового набору та безпечна реформуляція або реферування повідомлення. Третій рівень виконує функцію інтегральної підтримки прийняття рішень: він агрегує сигнали попередніх підсистем, ураховує новизну, масштаб поширення та невизначеність висновку і формує пріоритет випадку для подальшого опрацювання аналітиком. Запропонована методологія побудови системи базується на таких принципах: – модульне розділення на підсистеми аналізу повідомлення, добування свідчень, безпечної реформуляції та прийняття рішень; – доступність і інтерпретованість проміжних результатів кожного етапу; – орієнтація на українськомовні новинні та соціально-медійні дані; – можливість масштабування системи за рахунок заміни або розширення окремих модулів; – придатність до інтеграції у реальні програмні контури інформаційноаналітичного призначення. У другому розділі досліджено нейромережеві методи аналізу окремого повідомлення. Особливу увагу приділено проблемі використання великих мовних моделей у дискримінативних задачах. У роботі запропоновано підхід до побудови двоспрямованого енкодера на базі декодерної великої мовної моделі Gemma2-27B. Для цього каузальну маску замінено на повну двоспрямовану, а модель доадаптовано за схемою masked language modeling на змішаному українсько- та російськомовному новинному корпусі. Такий підхід дозволив сформувати модель biGemma2-27B, орієнтовану на українськомовні задачі класифікації та локалізації маніпулятивного контенту. Експериментальні результати показали, що biGemma2-27B продемонструвала найкращі результати серед розглянутих відкритих моделей. У третьому розділі досліджено доказово-орієнтовану обробку медіатекстів. Для цього розроблено синтетичний корпус EviManip-UA, призначений для моделювання задачі відновлення нейтрального першотексту за маніпулятивним переписуванням, добування свідчень і безпечної реформуляції. Такий корпус виконує подвійну роль: з одного боку, він є інструментом для побудови доказово-орієнтованих сценаріїв, а з іншого – слугує засобом оцінювання практичної придатності доказово-орієнтованих методів. Для масштабованого добування документів запропоновано використання напівкерованого інвертованого файлового індексу SS-IVF. На відміну від стандартного IVF, у запропонованому підході структуру розбиття простору узгоджено з навчальними сигналами задачі, що дозволяє краще зберігати релевантні документи в близьких кластерах. У роботі також розроблено підхід до формування доказового набору, який не зводиться до простого списку найближчих документів. Свідчення поділяються на такі групи: фрагменти, що підтверджують гіпотезу; фрагменти, що суперечать їй; подібні історичні прецеденти; та лексично схожі, але змістово нерелевантні приклади. Така типізація дає змогу перейти від суто пошукового модуля до засобу підтримки аналітичного міркування, де важливими є не лише підтвердження, а й суперечності, альтернативи та межі достовірності автоматичного висновку. У четвертому розділі наведено результати експериментальних досліджень та запропоновано архітектуру програмної реалізації системи. Показано, що всі ключові компоненти – модуль аналізу повідомлення, контур видобування, шар реформуляції – можуть бути інтегровані в єдину прикладну систему. Визначено основні функціональні підсистеми, потоки даних між ними, модель збереження результатів і принципи побудови інтерфейсу аналітика. Запропонований інтерфейс орієнтовано на практичну роботу з повідомленнями, доказовими наборами, безпечними викладами та фінальним рішенням щодо кейсу. Підсумовуючи, наукова новизна дисертаційної роботи полягає в наступному. Вперше запропоновано формалізацію інтелектуальної системи підтримки прийняття рішень для виявлення медіаманіпуляцій як трирівневої доказовоорієнтованої системи, що поєднує аналіз окремого повідомлення, добування та типізацію свідчень, безпечну реформуляцію і шар інтегрального прийняття рішень для аналітика. Вперше запропоновано підхід до побудови українсько-орієнтованого двоспрямованого енкодера на базі декодерної великої мовної моделі шляхом заміни каузальної маски на повну двоспрямовану та подальшої доменної MLMадаптації на змішаному українсько- та російськомовному новинному корпусі, що підвищує придатність моделі до дискримінативних задач аналізу тексту. Вперше адаптовано та обґрунтовано використання GRPO для задач безпечної реформуляції та реферування українськомовних медіатекстів із комбінованою функцією винагороди, яка одночасно враховує нетоксичність, покриття свідчень, нейтральність викладу та контрольовану довжину вихідного тексту. Удосконалено метод напівкерованого інвертованого файлового індексу SS-IVF для задачі добування свідчень і відновлення нейтрального першотексту за маніпулятивним переписуванням, що забезпечує кращий компроміс між швидкодією та якістю видобування у шумному інформаційному середовищі. Набули подальшого розвитку методи доменної адаптації нейромережевих моделей до українських новинних і соціально-медійних даних за рахунок поєднання двомовного новинного донавчання, цільового прикладного донавчання та доказово-орієнтованого оцінювання в межах єдиної системної постановки. Практичне значення одержаних результатів полягає в тому, що запропоновані методи, моделі та архітектурні рішення можуть бути безпосередньо використані для побудови прикладних систем моніторингу й аналізу медіаконтенту, орієнтованих на підтримку роботи аналітиків, дослідників інформаційних операцій, фахівців з OSINT, фактчекінгу та медіабезпеки. Усі теоретичні та практичні результати дисертаційної роботи опубліковано у фахових вітчизняних і закордонних наукових виданнях та апробовано на міжнародних наукових конференціях. За матеріалами дисертації опубліковано 5 наукових праць, з яких 2 – статті у журналах, що індексується у міжнародній наукометричній базі Scopus, 1 – стаття у фаховому виданні категорії Б, 2 – тези міжнародних наукових конференцій.
dc.description.abstractotherBazdyrev A. Neural Network-Based Intelligent Decision Support System for Media Manipulation Detection. – Qualifying scientific work submitted as a manuscript. Dissertation for the degree of Doctor of Philosophy in specialty 124 System Analysis (12 – Information Technology). – National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute”, Kyiv, 2026. In the modern information environment, media manipulation is one of the key instruments of informational influence, which has become especially acute in the context of the Russian-Ukrainian war, the massive dissemination of messages through social networks, messengers, and digital media, as well as the highspeed replication of manipulative narratives in different textual variations. For the Ukrainian information space, this problem is not only scientific but also directly practical, since timely detection of manipulative influence is important for security analytics, OSINT research, monitoring of information operations, fact-checking, and media literacy. At the same time, in the practical work of an analyst, it is not enough merely to determine whether a particular message is manipulative. It is also necessary to understand which specific technique has been used, which fragments of the message are trigger elements, which external evidence confirms or refutes the system’s conclusion, how to safely formulate a neutral account, and which cases should be considered first under limited attention resources. In this regard, the dissertation considers the problem of constructing an intelligent decision support system for detecting media manipulation. The proposed formulation combines methods of systems analysis, neural natural language processing, information retrieval, safe text generation, and practical reranking of cases. The central idea of the work is the transition from separate models to a coherent evidence-oriented architecture in which the outputs of neural modules are interpreted as elements of an analytical pipeline that culminates in practically applicable decision support. In the first chapter, a three-level problem formulation is proposed. The first level is responsible for the analysis of an individual message and includes determining the risk of manipulative influence, multi-label classification of techniques, and localization of trigger fragments. The second level implements an evidence-oriented pipeline within which evidence retrieval, reranking of retrieved candidates, typing of the evidence set, and safe reformulation or summarization of the message are performed. The third level performs the function of integral decision support: it aggregates signals from previous subsystems, takes into account novelty, propagation scale, and uncertainty of the conclusion, and forms a case priority for further analyst review. The proposed methodology for constructing the system is based on the following principles: – modular separation into subsystems for message analysis, evidence retrieval, safe reformulation, and decision making; – availability and interpretability of intermediate results at each stage; – orientation toward Ukrainian-language news and social media data; – the possibility of scaling the system by replacing or extending individual modules; – suitability for integration into real information-analytical software pipelines. The second chapter investigates neural methods for the analysis of an individual message. Particular attention is paid to the problem of using large language models in discriminative tasks. The dissertation proposes an approach to constructing a bidirectional encoder based on the decoder-only large language model Gemma2-27B. For this purpose, the causal mask is replaced with a full bidirectional one, and the model is further adapted using masked language modeling on a mixed Ukrainian- and Russian-language news corpus. This approach made it possible to construct the model biGemma2-27B, oriented toward Ukrainian-language tasks of classification and localization of manipulative content. Experimental results showed that biGemma2-27B demonstrated the best results among the considered open models. The third chapter investigates evidence-oriented processing of media texts. For this purpose, the synthetic corpus EviManip-UA was developed, intended for modeling the task of recovering a neutral source text from manipulative rewriting, evidence retrieval, and safe reformulation. This corpus performs a dual role: on the one hand, it is a tool for constructing evidence-based scenarios, and on the other hand, it serves as a means of evaluating the practical applicability of evidence-oriented methods. For scalable document retrieval, the use of the semi-supervised inverted file index SS-IVF is proposed. Unlike standard IVF, in the proposed approach the partitioning structure of the space is aligned with the learning signals of the task, which makes it possible to preserve relevant documents better in nearby clusters. The dissertation also develops an approach to constructing an evidence set that is not reduced to a simple list of nearest documents. Evidence is divided into the following groups: fragments that support the hypothesis; fragments that contradict it; similar historical precedents; and lexically similar but semantically irrelevant examples. Such typing makes it possible to move from a purely retrieval-oriented module to a tool that supports analytical reasoning, where not only confirmations but also contradictions, alternatives, and the limits of reliability of the automatic conclusion are important. The fourth chapter presents the results of experimental studies and proposes the architecture of the software implementation of the system. It is shown that all key components – the message analysis module, the retrieval pipeline, and the reformulation layer – can be integrated into a unified applied system. The main functional subsystems, data flows between them, the result storage model, and the principles of constructing the analyst interface are determined. The proposed interface is oriented toward practical work with messages, evidence sets, safe reformulations, and the final decision on a case. In summary, the scientific novelty of the dissertation lies in the following. For the first time, the formalization of an intelligent decision support system for detecting media manipulation as a three-level evidence-oriented system combining analysis of an individual message, retrieval and typing of evidence, safe reformulation, and a layer of integral decision making for the analyst has been proposed. For the first time, an approach to constructing a Ukrainian-oriented bidirectional encoder based on a decoder-only large language model has been proposed by replacing the causal mask with a full bidirectional one and subsequently performing domain-specific MLM adaptation on a mixed Ukrainian- and Russianlanguage news corpus, which increases the suitability of the model for discriminative text analysis tasks. For the first time, the use of GRPO has been adapted and substantiated for the tasks of safe reformulation and summarization of Ukrainian-language media texts with a combined reward function that simultaneously takes into account non-toxicity, evidence coverage, neutrality of presentation, and controlled output length. The method of the semi-supervised inverted file index SS-IVF has been improved for the task of evidence retrieval and recovery of a neutral source text from manipulative rewriting, which provides a better compromise between speed and retrieval quality in a noisy information environment. Methods of domain adaptation of neural network models to Ukrainian news and social media data have been further developed by combining bilingual news adaptation, task-specific applied fine-tuning, and evidence-oriented evaluation within a unified system formulation. The practical significance of the obtained results lies in the fact that the proposed methods, models, and architectural solutions can be directly used for building applied systems for monitoring and analyzing media content oriented toward supporting the work of analysts, researchers of information operations, specialists in OSINT, fact-checking, and media security. All theoretical and practical results of the dissertation have been published in recognized national and international scientific publications and presented at international scientific conferences. Based on the dissertation materials, 5 scientific works have been published, including 1 article in a journal indexed in the international scientific database Scopus, 1 article in a professional journal of category B, 1 single-author book chapter indexed in the international scientific database Scopus, and 2 proceedings of international scientific conferences.
dc.format.extent160 с.
dc.identifier.citationБаздирев, А. А. Інтелектуальна система підтримки прийняття рішень на основі нейронних мереж для виявлення медіаманіпуляцій : дис. … д-ра філософії : 124 Системний аналіз / Баздирев Антон Андрійович. - Київ, 2026. - 160 с.
dc.identifier.urihttps://ela.kpi.ua/handle/123456789/82175
dc.language.isouk
dc.publisherКПІ ім. Ігоря Сікорського
dc.publisher.placeКиїв
dc.subjectвеликі мовні моделі
dc.subjectобробка природної мови
dc.subjectтекстова аналітика
dc.subjectпошук релевантної інформації
dc.subjectоптимізація
dc.subjectмедіаманіпуляції
dc.subjectштучний інтелект
dc.subjectмашинне навчання
dc.subjectглибоке навчання
dc.subjectнейронні мережі
dc.subjectтрансформери
dc.subjectоцінка ризиків
dc.subjectгібридна система
dc.subjectсистема підтримки прийняття рішень
dc.subjectlarge language models
dc.subjectnatural language processing
dc.subjecttext analysis
dc.subjectretrieval
dc.subjectoptimization
dc.subjectmedia manipulation
dc.subjectartificial intelligence
dc.subjectmachine learning
dc.subjectdeep learning
dc.subjectneural networks
dc.subjecttransformers
dc.subjectrisk estimation
dc.subjecthybrid system
dc.subjectdecision support system
dc.subject.udc004.891.2
dc.titleІнтелектуальна система підтримки прийняття рішень на основі нейронних мереж для виявлення медіаманіпуляцій
dc.title.alternativeNeural Network-Based Intelligent Decision Support System for Media Manipulation Detection
dc.typeThesis Doctoral

Файли

Контейнер файлів
Зараз показуємо 1 - 1 з 1
Вантажиться...
Ескіз
Назва:
Bazdyrev_dys.pdf
Розмір:
9.22 MB
Формат:
Adobe Portable Document Format
Ліцензійна угода
Зараз показуємо 1 - 1 з 1
Ескіз недоступний
Назва:
license.txt
Розмір:
8.98 KB
Формат:
Item-specific license agreed upon to submission
Опис: