Artificial Intelligence · 01.09.2026, 11:48 UTC
Interpretable Predictability-Based AI Text Detection: A Replication Study
| Schweregrad | info |
|---|---|
| Kategorie | Artificial Intelligence |
| Quelle | arXiv cs.AI ↗ |
| Veröffentlicht | 01.09.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
arXiv:2603.15034v2 Announce Type: replace-cross Abstract: This paper replicates and extends the system used in the AuTexTification shared task for authorship attribution of machine-generated texts. Exact replication was not possible because of differences in data splits, model availability, and implementation details, which we document as a case study in reproducibility. We tested newer multilingual language models (mDeBERTa-v3-base, Qwen, mGPT) and added 26 document-level stylometric features, using ablation, permutation importance, and SHAP analysis to assess feature influence. A single shared configuration was applied to both English and Spanish across Subtask 1 and Subtask 2. Averaged over three random seeds, the shared multilingual configuration performs comparably to or better than the language-specific baseline, with the clearest gains on model attribution (Subtask 2). The additional stylometric features yield small improvements, led by lexical diversity, but their contribution falls within seed variance once predictability-based probabilities are included, which remain the dominant signal. The study also shows that clear documentation is important for reliable replication and fair comparison of systems.
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info SPADE: Self-Play in Adaptive Synthetic Executable Environments
- info Decomposing Wrong-Consensus Agreement in LLM Self-Consistency
- info Modeling the Structure of Human Behavior with AI Prompt Vectors
- info Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models