Data Engineer | Python • SQL • Azure • AWS • Databricks • PySpark • dbt • RAG
I build reliable cloud data pipelines, lakehouse and warehouse solutions, analytics-ready datasets, and AI/RAG-enabled data applications with strong documentation, testing, and business-focused reporting.
I’m Rihua Van Steenburgh, a Data Engineer focused on building end-to-end data pipelines, cloud data platforms, analytics-ready datasets, and AI/RAG-enabled data applications.
I am pursuing a Master of Science in Data Analytics at Middle Georgia State University, expected in December 2026. My background in data analytics, web development, and cloud data engineering helps me connect technical pipeline design with business reporting needs.
My portfolio features hands-on projects across Azure, AWS, Databricks, PySpark, SQL, dbt, Redshift, Docker,PostgreSQL/pgvector, and RAG. I focus on building projects that are reproducible, well-documented, tested, and easy for recruiters and teams to evaluate.
-
Cloud Data Engineering: Azure Data Factory, ADLS Gen2, Databricks, PySpark, Delta Lake, AWS S3, Redshift Serverless
-
Data Pipelines: API ingestion, batch processing, orchestration, scheduling, and workflow automation
-
Data Modeling: Bronze/Silver/Gold layers, dimensional modeling, star schema, fact/dimension tables, and analytics marts
-
Data Quality & Reliability: validation checks, testing, clean transformations, CI workflows, runbooks, and documentation
-
AI Data Engineering: document ingestion, chunking, embeddings, PostgreSQL/pgvector, vector search, RAG, and cited answers
-
BI & Communication: Power BI dashboards, KPI reporting, architecture diagrams, and clear technical storytelling
-
Languages: Python, SQL, R, JavaScript, TypeScript
-
Cloud & Data Platforms: Azure Data Factory, ADLS Gen2, Databricks, AWS S3, Redshift Serverless, ECS/Fargate, EventBridge, CloudWatch
-
Processing & Modeling: PySpark, Delta Lake, dbt, PostgreSQL, dimensional modeling, star schema
-
AI / RAG: document ingestion, chunking, embeddings, PostgreSQL/pgvector, vector search, retrieval-augmented generation, cited answers
-
Orchestration & CI/CD: Airflow, GitHub Actions, workflow automation, validation checks
-
BI & Tools: Power BI, DAX, Docker, Git/GitHub, Jupyter Notebook, VS Code
-
Web & APIs: REST APIs, JSON, CSV, WordPress
Azure data engineering lakehouse for NYC 311 operational analytics. Ingests service request data using Azure Data Factory, lands raw data in ADLS Gen2, and processes Bronze, Silver, and Gold Delta Lake layers in Databricks with Python/PySpark and SQL. Produces analytics-ready dimensions, fact tables, validation checks, Gold marts, runbooks, architecture notes, and cloud execution proof.
Tech: Python, SQL, Azure Data Factory, ADLS Gen2, Databricks, PySpark, Delta Lake, GitHub Actions, Power BI-ready outputs
AWS batch data pipeline for external flight fare API data. Runs Python container jobs with Docker and ECS/Fargate, lands raw data in S3 Bronze, loads Redshift Serverless, and transforms data with SQL-based dbt staging models, dimensional marts, and tests. Documents CloudWatch proof logs, CI checks, architecture diagrams, runbooks, and cost/secret safety notes
Tech: Python, SQL, Docker, ECS/Fargate, EventBridge Scheduler, S3, Redshift Serverless, dbt, CloudWatch, GitHub Actions
Local AI data engineering and RAG copilot for NYC 311 documentation. Ingests curated documents, chunks source text, stores embeddings in PostgreSQL/pgvector, retrieves cited context, and presents grounded answers in a Streamlit UI with source citations, retrieved chunk previews, Dockerized pgvector setup, pytest coverage, GitHub Actions CI, and an 18-question evaluation set.
Tech: Python, PostgreSQL, pgvector, Streamlit, Docker, embeddings, vector search, RAG, pytest, GitHub Actions
- Improving cloud data engineering portfolio projects with Azure, AWS, Databricks, dbt, CI, tests, and runbooks
- Expanding AI data engineering work with RAG, PostgreSQL/pgvector, cited answers, and evaluation sets
- Building recruiter-ready project documentation, architecture diagrams, and case studies
Thanks for stopping by! ✨


