saeid@brussels:~$ whoami

Hi, I'm Saeid Shahriari

Data Engineer·MLOps / MLSecOps·Brussels

I design and build reliable batch and streaming data pipelines that are security-first, from raw source to a live dashboard. I care about data that is trustworthy, observable, and safe.

streaming pipeline · live
source API CDC Kafka Flink Spark lakehouse warehouse dashboard
pipeline healthy · 0 failed tasks · data flowing
0
GitHub repositories
0
Years engineering
0
Certifications earned
0
Tools & platforms
Saeid Shahriari
Open to work · Brussels
// about

From critical systems to data pipelines

I am a Data Engineer based in Brussels. Since September 2025 I have worked independently on data engineering and MLOps projects, after deepening my skills in an MLOps-focused data engineering program. My focus is reliable pipelines, data quality, and security.

I bring more than a decade of engineering discipline, from critical medical systems to modern data platforms, where reliability and compliance are not optional. I specialize in real-time streaming (Kafka / Flink), lakehouse architectures, and data security (MLSecOps), because a pipeline is only useful when its data can be trusted.

StreamingLakehouseMLOpsData SecurityObservabilityGDPR
Let's connect
// tech stack

Tools I build with

The technologies I use to move data reliably and safely, from source to insight.

Languages

Python (Pandas, PySpark, PyFlink) · SQL (PostgreSQL, ClickHouse) · Java · Bash

Streaming & Processing

Apache Kafka · Apache Flink · Spark Structured Streaming · Debezium (CDC)

Orchestration

Apache Airflow · Docker · Kubernetes · Jenkins · GitHub Actions (CI/CD)

Lakehouse & Warehouse

Databricks (Delta Lake, DLT) · Snowflake · dbt · Apache Spark · Parquet

Cloud & Platforms

AWS (Glue, Athena, Lambda) · GCP (BigQuery)

Security & MLOps

HashiCorp Vault · Prompt-injection scanning · MLflow · Grafana · Prometheus

// what I build

From raw data to production

Systems I design, build, secure, and monitor end to end.

Data Pipelines

Reliable batch and streaming pipelines that move data from source to storage with Kafka, Flink & Spark.

Lakehouse & Warehousing

Medallion (Bronze → Silver → Gold) models on Delta Lake, Snowflake and dbt. Fast, clean, and query-ready.

MLOps & Model Serving

Model APIs with monitoring, MLflow tracking, drift detection and retraining built in from day one.

Data Security & Governance

Secrets management, input scanning, anonymization & GDPR-conscious design. Security is a first layer, not an afterthought.

Observability

Grafana & Prometheus dashboards and alerts so issues surface early and pipelines stay healthy 24/7.

Data Visualization

Clear dashboards in Streamlit, Tableau & Power BI that turn pipelines into decisions.

// selected projects

Things I've built

Real, end-to-end projects across streaming, lakehouse, MLOps and data security. Every one is on GitHub.

Secure Lakehouse · MLSecOps

ETRM Data Platform

An automated, secured Energy Trading & Risk Management lakehouse. It ingests real European electricity prices, defends the pipeline across seven security layers, transforms everything with Spark (Bronze → Silver → Gold), flags suspicious trades with an Isolation Forest surveillance model, and serves a live dashboard plus a read-only AI agent you can ask in plain language.

AirflowSparkPostgreSQLVaultMLflowDuckDBStreamlitDocker
View project
Risk-first Data Engineering

Belgian Deeds Lakehouse

A risk-first pipeline for Belgian Gazette incorporation deeds. It ingests real deed PDFs, runs OCR where needed, scans untrusted input for malicious content & prompt injection, extracts structured data with Gemini, validates it with Pydantic, stores it in PostgreSQL, and serves it through a FastAPI REST API and a read-only AI agent.

PythonFastAPIPostgreSQLGeminiTesseract OCRPydanticDocker
View project
Streaming Platform

Real-Time Food Delivery Platform

Production-style CDC pipeline: PostgreSQL changes flow into Kafka, Flink SQL materializes streaming views, and online outputs land in OpenSearch and Redis. Built with Docker Compose, repeatable connector registration, and verification scripts.

KafkaFlink SQLPostgreSQLDebeziumRedisOpenSearch
View project
Analytics Engineering

ELT Pipeline · dbt · Snowflake · Airflow

Analytics engineering workflow using Snowflake sample data, dbt models and tests, and Airflow via Astronomer Cosmos. Produces a clean fact table with automated data-quality checks and orchestration around transformation dependencies.

dbtSnowflakeAirflowCosmosSQL
View project
MLOps / Model Serving

Real-Time Fraud Detection API

A FastAPI fraud-scoring service with health checks, feedback capture, and monitoring endpoints, integrated with Prometheus, Grafana, and MLflow. Designed as a practical starting point for near real-time model operations.

FastAPIPrometheusGrafanaMLflowDocker
View project
Privacy Engineering

PostgreSQL Anonymizer Streaming

A privacy-by-design demo built fully inside PostgreSQL. It combines dynamic masking, trigger-based anonymized streaming into a sanitized schema, and static anonymization for safe data exports in GDPR-conscious environments.

PostgreSQLTriggersHMAC TokenizationDocker
View project
MLOps · Drift & Retraining

Real-Time Fraud Detection Pipeline

A fully containerized fraud-detection pipeline: a Java producer streams synthetic transactions into Kafka, Spark Structured Streaming computes features in real time, PostgreSQL stores the scores, and Airflow orchestrates drift detection and model retraining.

KafkaSpark StreamingPostgreSQLAirflowJava
View project
// experience

My journey

More than a decade of engineering, from critical medical systems to modern data platforms.

Data & Engineering

Data Engineer · Self-employed

Sep 2025 - Present · Belgium · Remote

Working independently on data engineering and MLOps projects after an MLOps-focused data engineering program. Focused on reliable pipelines, data quality, and security, across the streaming, lakehouse and MLOps projects shown above.

Data Visualization (Student Job)

Orange Business · Jul 2025 - Aug 2025 · Brussels · On-site

Used Google Cloud Platform and Power BI to visualize complex datasets, supported ETL processes, and produced clear technical reports for business stakeholders.

Data Engineer (Internship)

Data Engineering Hub · Nov 2024 - Mar 2025 · Remote

Built a fraud detection system with Apache Kafka, Apache Spark, PostgreSQL, MySQL and Grafana for an organization with multiple offices and databases, working with cross-functional teams. Also completed Data Engineering with AWS (CloudFormation).

Data Science with a Flavor of Data Engineering (Internship)

Data Science School · Nov 2024 - Feb 2025 · Remote

Built an Employee Attendance System, taking a machine-learning model from concept to cloud. Deployed a Streamlit app to Hugging Face via GitHub Actions.

Medical Engineering Background

Medical Device Service Engineer

Sina Parto Jam Biomedical & Engineering Co. · Apr 2019 - Dec 2021 · Tehran, Iran

Installed, modified and repaired medical imaging equipment (including digital X-ray) at customer sites. Analyzed mechanical, electrical and software failures, ensured FDA regulatory compliance, and grew support-service revenue by 68%.

Radiology Technician

Bu Ali Sina Hospital · Aug 2014 - Feb 2017 · Hamadan, Iran

Performed quality imaging studies, operated imaging equipment safely across sites, and ensured patient safety through quality assurance and quality control.

Medical Physicist

Mahdieh Diagnostic & Treatment Center · Jul 2012 - Sep 2012 · Iran

Supported safe delivery of radiation therapy, treatment planning and acceptance testing, and ran quality-assurance surveys (image quality, calibration and safety) on imaging systems.

// education

Academic background

Applied Computer Science, VUB

Vrije Universiteit Brussel (VUB) · Brussels, Belgium

Graduate coursework in Data Engineering, Machine Learning, Big Data and Distributed Systems, including Big Data Processing, AI Techniques, Advanced Databases, Cloud & Distributed Systems, and Advanced IT Networks.

// certifications

Courses & certificates

Selected training that supports my data engineering path.

Data Engineering: Local Dev to Server Deployment

Data Engineering School · Mar 2025

CI/CD, PostgreSQL, Git, Kubernetes, Airflow, Docker, AWS, Kafka, ETL.

Hands-On Essentials: Data Warehousing

Snowflake · Jun 2025

Data warehousing fundamentals and best practices on Snowflake.

Introduction to dbt

DataCamp · Jun 2025

Analytics engineering with dbt for modular SQL transformations and testing.

Data Ingestion with Delta Lake

Databricks · Apr 2025

Reliable ingestion pipelines with Delta Lake and the Lakehouse.

Data Management & Governance

Databricks · Apr 2025

Governance, access control and quality on the Databricks Lakehouse.

Foundational Cloud Practitioner

Maktabkhooneh · Mar 2025

Core cloud concepts, security and cost-awareness.

Machine Learning

Coursera · Jun 2021

Classical ML algorithms, regression, neural networks, model evaluation.

Advanced Python Programming

Maktabkhooneh · Apr 2021

Advanced Python for data processing, scripting and backend development.

// writing

Selected writing

Notes on data engineering, shared on LinkedIn.

28 Nov

Soft skills that really matter in tech

Read on LinkedIn
27 Nov

Data Engineering for Solution Architecture

Read on LinkedIn
04 Nov

ETL → ELT in 2025: what changed and what actually works

Read on LinkedIn
// contact

Let's build something reliable together

Open to Data Engineering & MLOps opportunities in Brussels and remote. Have a pipeline to build, a dataset to tame, or a role in mind? Send a message.

info@saeidshahriari.com

Privacy