I design and build reliable batch and streaming data pipelines that are security-first, from raw source to a live dashboard. I care about data that is trustworthy, observable, and safe.
I am a Data Engineer based in Brussels. Since September 2025 I have worked independently on data engineering and MLOps projects, after deepening my skills in an MLOps-focused data engineering program. My focus is reliable pipelines, data quality, and security.
I bring more than a decade of engineering discipline, from critical medical systems to modern data platforms, where reliability and compliance are not optional. I specialize in real-time streaming (Kafka / Flink), lakehouse architectures, and data security (MLSecOps), because a pipeline is only useful when its data can be trusted.
Let's connectThe technologies I use to move data reliably and safely, from source to insight.
Python (Pandas, PySpark, PyFlink) · SQL (PostgreSQL, ClickHouse) · Java · Bash
Apache Kafka · Apache Flink · Spark Structured Streaming · Debezium (CDC)
Apache Airflow · Docker · Kubernetes · Jenkins · GitHub Actions (CI/CD)
Databricks (Delta Lake, DLT) · Snowflake · dbt · Apache Spark · Parquet
AWS (Glue, Athena, Lambda) · GCP (BigQuery)
HashiCorp Vault · Prompt-injection scanning · MLflow · Grafana · Prometheus
Systems I design, build, secure, and monitor end to end.
Reliable batch and streaming pipelines that move data from source to storage with Kafka, Flink & Spark.
Medallion (Bronze → Silver → Gold) models on Delta Lake, Snowflake and dbt. Fast, clean, and query-ready.
Model APIs with monitoring, MLflow tracking, drift detection and retraining built in from day one.
Secrets management, input scanning, anonymization & GDPR-conscious design. Security is a first layer, not an afterthought.
Grafana & Prometheus dashboards and alerts so issues surface early and pipelines stay healthy 24/7.
Clear dashboards in Streamlit, Tableau & Power BI that turn pipelines into decisions.
Real, end-to-end projects across streaming, lakehouse, MLOps and data security. Every one is on GitHub.
An automated, secured Energy Trading & Risk Management lakehouse. It ingests real European electricity prices, defends the pipeline across seven security layers, transforms everything with Spark (Bronze → Silver → Gold), flags suspicious trades with an Isolation Forest surveillance model, and serves a live dashboard plus a read-only AI agent you can ask in plain language.
A risk-first pipeline for Belgian Gazette incorporation deeds. It ingests real deed PDFs, runs OCR where needed, scans untrusted input for malicious content & prompt injection, extracts structured data with Gemini, validates it with Pydantic, stores it in PostgreSQL, and serves it through a FastAPI REST API and a read-only AI agent.
Production-style CDC pipeline: PostgreSQL changes flow into Kafka, Flink SQL materializes streaming views, and online outputs land in OpenSearch and Redis. Built with Docker Compose, repeatable connector registration, and verification scripts.
Analytics engineering workflow using Snowflake sample data, dbt models and tests, and Airflow via Astronomer Cosmos. Produces a clean fact table with automated data-quality checks and orchestration around transformation dependencies.
A FastAPI fraud-scoring service with health checks, feedback capture, and monitoring endpoints, integrated with Prometheus, Grafana, and MLflow. Designed as a practical starting point for near real-time model operations.
A privacy-by-design demo built fully inside PostgreSQL. It combines dynamic masking, trigger-based anonymized streaming into a sanitized schema, and static anonymization for safe data exports in GDPR-conscious environments.
A fully containerized fraud-detection pipeline: a Java producer streams synthetic transactions into Kafka, Spark Structured Streaming computes features in real time, PostgreSQL stores the scores, and Airflow orchestrates drift detection and model retraining.
More than a decade of engineering, from critical medical systems to modern data platforms.
Data & Engineering
Working independently on data engineering and MLOps projects after an MLOps-focused data engineering program. Focused on reliable pipelines, data quality, and security, across the streaming, lakehouse and MLOps projects shown above.
Used Google Cloud Platform and Power BI to visualize complex datasets, supported ETL processes, and produced clear technical reports for business stakeholders.
Built a fraud detection system with Apache Kafka, Apache Spark, PostgreSQL, MySQL and Grafana for an organization with multiple offices and databases, working with cross-functional teams. Also completed Data Engineering with AWS (CloudFormation).
Built an Employee Attendance System, taking a machine-learning model from concept to cloud. Deployed a Streamlit app to Hugging Face via GitHub Actions.
Medical Engineering Background
Installed, modified and repaired medical imaging equipment (including digital X-ray) at customer sites. Analyzed mechanical, electrical and software failures, ensured FDA regulatory compliance, and grew support-service revenue by 68%.
Performed quality imaging studies, operated imaging equipment safely across sites, and ensured patient safety through quality assurance and quality control.
Supported safe delivery of radiation therapy, treatment planning and acceptance testing, and ran quality-assurance surveys (image quality, calibration and safety) on imaging systems.
Graduate coursework in Data Engineering, Machine Learning, Big Data and Distributed Systems, including Big Data Processing, AI Techniques, Advanced Databases, Cloud & Distributed Systems, and Advanced IT Networks.
Selected training that supports my data engineering path.
CI/CD, PostgreSQL, Git, Kubernetes, Airflow, Docker, AWS, Kafka, ETL.
Data warehousing fundamentals and best practices on Snowflake.
Analytics engineering with dbt for modular SQL transformations and testing.
Reliable ingestion pipelines with Delta Lake and the Lakehouse.
Governance, access control and quality on the Databricks Lakehouse.
Core cloud concepts, security and cost-awareness.
Classical ML algorithms, regression, neural networks, model evaluation.
Advanced Python for data processing, scripting and backend development.
Notes on data engineering, shared on LinkedIn.
Open to Data Engineering & MLOps opportunities in Brussels and remote. Have a pipeline to build, a dataset to tame, or a role in mind? Send a message.
saeidshahriari1@gmail.com