A transparent ATS simulator built with PySpark, MongoDB, and Python to process resumes at scale, match candidates to jobs, and analyze real-time screening behavior through batch and streaming pipelines.
This project simulates a transparent Applicant Tracking System using PySpark, MongoDB, and Python. It processes resumes and job postings at scale, generates match scores with NLP techniques, stores results in MongoDB, and demonstrates both batch and streaming analytics for real-time screening behavior.
Kaggle Datasets: Resume dataset (snehaanbhawal/resume-dataset): 2,484 resumes in CSV format (~54 MB), with an ID, plain‑text resume, HTML resume, and one of 24 categories.
Job postings dataset (princekhunt19/700-jobs-data-of-ai-and-data-fields-2025): 735 listings scraped for the search term “data scientist” (2025). Fields include title, company, description, location, and salary when available.