Improvement of the method for speaker-metadata conditioning in fake-news detection

Loading...
Thumbnail Image

Date

item.page.thesis.degree.name

item.page.thesis.degree.level

item.page.thesis.degree.discipline

item.page.thesis.degree.department

item.page.thesis.degree.grantor

item.page.thesis.degree.advisor

item.page.thesis.degree.committeeMember

Journal Title

Journal ISSN

Volume Title

Publisher

Національний технічний університет "Харківський політехнічний інститут"

Abstract

The LIAR and LIAR2 benchmarks are among the most widely used datasets for automatic fake-news detection, and the speaker credit-history counts they provide are the main feature that distinguishes metadata-based detectors from plain text classifiers. If the gains these counts produce arise from how the dataset was built rather than from a real signal, then many leaderboardscomparisons measure something that cannot exist when a detector is actually deployed, which makes the integrity of these counts an important problem. The object of researchis the process of conditioning a transformer text encoder on speaker metadata for fake-news detection on the LIAR-family benchmarks. The subject of the researchis the temporal leakage contained in the pre-aggregated credit-history counts and the methods used to remove it before the features reach a model. The purpose of this paperis to improve the method of speaker-metadata conditioning by using credit-history features that satisfy the principle of temporal honesty, and to assess how much of the apparent metadata gain is leakage rather than usable signal. Research methods.A DeBERTa-v3 encoder is conditioned on the metadata through Feature-wise Linear Modulation under three regimes that differ only in how the counts are computed: leaky, as shipped; split-safe, using the standard remove-self correction; and temporal-honest, using only counts from a speaker's earlier statements.
Набори даних LIAR та LIAR2 є одними з найпоширеніших для автоматичного виявлення фейкових новин, а показники кредитної історії мовців, що містяться в них, є головною особливістю, яка відрізняє детектори на основі метаданих від класифікаторів простого тексту. Це робить ці підрахунки вартими ретельного вивчення: якщо отримані ними переваги походять від того, як був побудований набір даних, а не від реального сигналу, то багато порівнянь у рейтингах вимірюють те, чого ніколи не існувало б при фактичному використанні детектора. Ми зосереджуємося на часовій витоку, прихованій у цих попередньо агрегованих підрахунках, та на способах, якими люди намагаються її усунути, перш ніж ознаки потрапляють до моделі. Ми обумовлюємо кодер DeBERTa-v3 на метаданих за допомогою лінійної модуляції за ознаками (Feature-wise Linear Modulation) у трьох режимах, що відрізняються лише тим, як обчислюються підрахунки: «з витоком» (leaky), як у стандартній версії; «безпечний для розділення» (split-safe), з використанням стандартної корекції «видалення себе»; та «часово чесний» (temporal-honest), з використанням лише підрахунків із попередніх висловлювань мовця.

Description

Citation

Datsenko S., Kuchuk H., Piterska V., Omarov S. Improvement of the method for speaker-metadata conditioning in fake-news detection. Сучасні інформаційні системи. 2026. Т. 10, № 3. С. 97-103.

Endorsement

Review

Supplemented By

Referenced By