This repository contains the code and analysis notebooks accompanying the paper "Propaganda by Prompt: Tracing Hidden Linguistic Strategies in Large Language Models."
- ai_propaganda.db: SQLite database containing the article dataset and LIWC analysis results
- Must be downloaded from Harvard Dataverse (linked below)
- Contains both human-written and AI-generated articles
- Includes LIWC feature analysis for all articles
-
lgbm_train_models.py: Training script for LightGBM propaganda detection models
- Trains models for different topics and GPT versions
- Implements feature selection and normalization
- Produces model performance metrics
-
hyperparameter_grid_search.py: Hyperparameter optimization script for propaganda detection models
- Performs grid search on hyperparameters for Shallow & Deep Neural Networks, SVM, and LightGBM models
- Model_Analysis_Notebook.ipynb: Primary analysis notebook for model interpretation
- Loads existing models and queries the database
- Performs SHAP (SHapley Additive exPlanations) analysis
- Generates visualizations of feature importance
- Examines feature interactions and dependencies
The models/ directory contains trained LightGBM models (.pkl files) for:
- Combined analysis across all topics
- Topic-specific models (climate, COVID-19, Capitol riot, LGBT)
- Variants with and without punctuation features
- Different GPT versions (GPT-3.5, GPT-4o, GPT-4.1)
Recommended Python Version: 3.11.6
Required Python packages are listed in requirements.txt. Install using:
pip install -r requirements.txt- Download the
ai_propaganda.dbfile from Harvard Dataverse: https://doi.org/10.7910/DVN/PENY1S - Install required packages
- Run analysis notebooks for model interpretation
- Use training scripts to reproduce model results
If you use this code or dataset in your research, please cite:
Barfar, A., & Sommerfeldt, L. (2026). Propaganda by prompt: Tracing hidden linguistic strategies in large language models. Information Processing & Management, 63(2), 104403.
For questions about the code or data, please contact the authors through the paper's corresponding author.