I engineer automated Python ETL pipelines, advanced text-sanitization workflows, and data pipelines for commercial analytics and bioinformatics.
- Languages: Python (Pandas, NumPy, Regex, Biopython), Bash/Linux Terminal
- Data Engineering: Automated ETL Pipelines, Data Quality Assurance, Exploratory Data Analysis (EDA)
- SARS-CoV-2 Genomic Architecture Analysis: Custom sequence analysis pipeline built in Python utilizing string manipulation to parse genomic alignments, identify mutations, and trace viral structural transformations.
- Geo Micro-Array Dataset Exploratory Analysis: High-dimensional EDA and text-cleaning pipelines optimized to extract gene expression arrays and sanitize messy index noise.
- Automated Forex Trend & Volatility Analyzer: High-velocity Python engine that ingests live market parameters, processes asset volatility scripts, and computes clean trend signals.
- Telco Customer Churn Pipeline Optimization: End-to-end data processing workflow mapping messy customer profiles, treating null variables, and engineering features for predictive data models.
- Airbnb NYC 2019 Structural EDA: Complete geospatial and pricing data sanitation layout, parsing erratic strings and mapping regional inventory parameters cleanly.
- Titanic Predictive Feature Engineering Lab: Classic data quality sandbox focused on mathematical imputation, out-of-bounds drop constraints, and data type standardizations.
- 💼 Available for: Remote Engineering Contracts, White-Label Agency Subcontracting, and Research Placements.
- 📬 Inbox Channel: www.linkedin.com/in/ndubuisi-akin-00b243209