KH Coder: for Quantitative Content Analysis or Text Mining
-
Updated
Aug 22, 2026 - Perl
KH Coder: for Quantitative Content Analysis or Text Mining
Significance testing for collocates.
A Python toolkit for processing and analyzing the 21st Century Sejong Corpus (National Institute of Korean Language) | 21세기 세종계획 말뭉치 분석 도구
I use various techniques for analyzing the Stanford Congressional Records. Specifically, we will be looking at
An R library for displaying "key word in context" (kwic)
Greek Stock Exchange Corpus: σώμα ενωσιακών κειμένων χρηματιστηριακού δικαίου (EL-FR) και εργαλείο διερεύνησης
A pipeline that transcribes audio to text using OpenAI's Whisper, extracts keywords with NLTK, and enables contextual keyword search via a Streamlit interface.
The KWIC index system accepts an ordered set of lines, each line is an ordered set of words, and each word is an ordered set of characters. Any line may be “circularly shifted” by repeatedly removing the first word and appending it at the end of the line. The KWIC index system outputs a listing of all circular shifts of all lines in alphabetical…
Extracts sentence, paragraph, and metadata for every keyword hit across NewsBank PDF exports, output to spreadsheet.
A software product line of Key Word in Context (KWIC).
Multilingual corpus analysis for EPUB and text files, with a .NET CLI, Avalonia desktop app, SQLite storage, KWIC, n-grams, collocations and run comparison.
To associate your repository with the kwic topic, visit your repo's landing page and select "manage topics."