I'm an early-career data engineer from Santo Domingo, Dominican Republic, building my way into data through real projects.
- 🎓 MSc in Data Science & Business Analytics (IMF Smart Education / UCAV, co-developed with Indra–Minsait)
- 📚 Currently studying Software Engineering to pair data skills with solid engineering practices
- 💼 ~6 years in customer service across regulated health & finance sectors — real domain knowledge and the habit of adapting fast to complex business rules
- 🌎 Spanish (native) · English (C1) · open to remote roles
- 🎯 Goal: build efficient, scalable and maintainable data pipelines — and become the best data engineer I can be
An intelligent asset-selection system for the US stock market. Built as my Master's thesis.
A multi-source ETL pipeline + ML system that scans ~3,600 US stocks using 34 features and 5 years of history, with strict point-in-time safety to avoid look-ahead bias.
- 🔗 Multi-source ingestion — SimFin, FRED, SEC EDGAR & Yahoo Finance merged into a single dataset
- 📦 Parquet storage with incremental per-ticker cache — daily screening dropped from ~60 min to ~5 min
- ⏱️ Point-in-time integrity — every source uses its real publication date
- ✅ CI/CD (GitHub Actions), 191 unit tests, MLflow tracking, drift monitoring
Python · pandas · scikit-learn · MLflow · Parquet · GitHub Actions · pytest
➡️ Repo coming soon — currently uploading the latest rebuild.
I use AI-assisted development as part of my workflow, and I'm open about it. My focus is on owning the decisions — architecture, data integrity, trade-offs — and understanding every part of what I ship well enough to explain and extend it. I treat AI as a power tool, not a substitute for understanding.
AWS Academy — ML & Cloud Foundations · IBM Data Science · Google Cybersecurity · Alura LATAM — Data Science & AI Agents
