Skip to content

Repository files navigation

Propaganda by Prompt: Repository

This repository contains the code and analysis notebooks accompanying the paper "Propaganda by Prompt: Tracing Hidden Linguistic Strategies in Large Language Models."

Repository Contents

Data Requirements

  • ai_propaganda.db: SQLite database containing the article dataset and LIWC analysis results
    • Must be downloaded from Harvard Dataverse (linked below)
    • Contains both human-written and AI-generated articles
    • Includes LIWC feature analysis for all articles

Python Scripts

  • lgbm_train_models.py: Training script for LightGBM propaganda detection models

    • Trains models for different topics and GPT versions
    • Implements feature selection and normalization
    • Produces model performance metrics
  • hyperparameter_grid_search.py: Hyperparameter optimization script for propaganda detection models

    • Performs grid search on hyperparameters for Shallow & Deep Neural Networks, SVM, and LightGBM models

Analysis Notebooks

  • Model_Analysis_Notebook.ipynb: Primary analysis notebook for model interpretation
    • Loads existing models and queries the database
    • Performs SHAP (SHapley Additive exPlanations) analysis
    • Generates visualizations of feature importance
    • Examines feature interactions and dependencies

Models Directory

The models/ directory contains trained LightGBM models (.pkl files) for:

  • Combined analysis across all topics
  • Topic-specific models (climate, COVID-19, Capitol riot, LGBT)
  • Variants with and without punctuation features
  • Different GPT versions (GPT-3.5, GPT-4o, GPT-4.1)

Requirements

Recommended Python Version: 3.11.6

Required Python packages are listed in requirements.txt. Install using:

pip install -r requirements.txt

Usage

  1. Download the ai_propaganda.db file from Harvard Dataverse: https://doi.org/10.7910/DVN/PENY1S
  2. Install required packages
  3. Run analysis notebooks for model interpretation
  4. Use training scripts to reproduce model results

Citation

If you use this code or dataset in your research, please cite:

Barfar, A., & Sommerfeldt, L. (2026). Propaganda by prompt: Tracing hidden linguistic strategies in large language models. Information Processing & Management, 63(2), 104403.

Contact

For questions about the code or data, please contact the authors through the paper's corresponding author.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages