00system.identity

DATA ENGINEER / RABAT, MOROCCO

Hamza
Bouali

I build data pipelines, streaming systems, and analytics platforms — on-premise, containerized, and cloud.

PIPELINE STATUSLIVE / 3Hz
BRBronze
→
SISilver
→
GOGold

All systems nominal last check 09:41:08

01projects.index

SELECTED SYSTEMS

Built to move
reliably.

A field log of pipelines, warehouses, and real-time systems. Select a node to inspect the architecture.

Microsoft Fabric / Medallion architecture900 tables / 20GB per client

Data-as-a-Service Pipeline

A repeatable data factory for ingesting large, multi-tenant on-premise estates into OneLake.

Microsoft FabricLakehouseDataflow Gen2Notebooks
Open case study ↘
02projects.case_studies
01CASE / FABRIC

Microsoft Fabric / Medallion architecture

Data-as-a-Service Pipeline

900 tables / 20GB per client

01 / Problem

Clients needed a dependable path from on-premise databases to analytics-ready data without rebuilding ingestion logic for every tenant.

02 / Architecture

01On-prem DBs
02JDBC
03Bronze
04Silver
05Gold

03 / Key decisions

  • Separated raw, refined, and serving layers to keep replay and quality checks explicit.
  • Used Dataflow Gen2 and notebooks for a mix of managed movement and custom transformations.
  • Designed refresh windows around full and incremental workloads.

04 / Outcome

40min full load · 5min incremental · 3 refreshes/day

05 / What I'd do differently

I would formalize tenant-level observability earlier, especially around freshness, failed partitions, and cost per refresh.

02CASE / BITCOIN

Streaming / per-batch model training

Bitcoin Real-Time ML Pipeline

Kafka → Spark → FastAPI

01 / Problem

A live market feed needed to move through ingestion, training, serving, and visualization without collapsing those concerns into one service.

02 / Architecture

01BTC-USD
02Kafka
03Spark
04FastAPI
05Streamlit

03 / Key decisions

  • Kept streaming computation in Spark Structured Streaming and model access behind a small API.
  • Used Docker Compose to make the complete multi-container system reproducible.
  • Exposed both real-time metrics and prediction endpoints for downstream consumers.

04 / Outcome

Live price stream · model metrics · prediction API

05 / What I'd do differently

I would add a stronger replay and evaluation harness so new model versions can be compared against identical historical windows.

03CASE / TAXI

Airflow / star-schema OLAP warehouse

NYC Green Taxi Data Pipeline

50GB / 71+ months

01 / Problem

A large historical dataset needed predictable monthly scheduling, validation, schema creation, and the ability to backfill without manual babysitting.

02 / Architecture

01CSV files
02Airflow
03Validate
04Warehouse
05OLAP

03 / Key decisions

  • Built chunked inserts and connection pooling around the warehouse boundary.
  • Made schema creation and validation part of the DAG rather than an external checklist.
  • Scheduled monthly loads while keeping backfill behavior explicit for 71+ months of history.

04 / Outcome

50GB loaded · 71+ months automated backfill

05 / What I'd do differently

I would expose per-month lineage and row-count drift as first-class run artifacts for faster operational review.

04CASE / BANKING

SQL Server / SSIS / Power BI

Banking BI System & Data Warehouse

DirectQuery + Import

01 / Problem

Banking reporting needed a stable analytical model and a BI layer that could balance freshness with dashboard performance.

02 / Architecture

01Source systems
02SSIS
03SQL Server
04Model
05Power BI

03 / Key decisions

  • Used a star schema to keep measures and dimensions legible to reporting users.
  • Combined DirectQuery and Import strategies according to dashboard needs.
  • Kept ETL responsibilities in SSIS and analytical presentation in Power BI.

04 / Outcome

Star schema · SSIS ETL · optimized Power BI dashboards

05 / What I'd do differently

I would define performance budgets per dashboard before tuning the storage mode, then make those trade-offs visible to stakeholders.

03about.credibility

OPERATING CONTEXT

Systems-minded.
Detail-aware.

My work sits between dependable infrastructure and the people who need to trust what comes out of it. I care about clear interfaces between stages, observable runs, and making the next change safer than the last.

EXPERIENCE LOG

02.2026—08.2026

Data Engineer

Veolia Software Solutions

On-premise multi-tenant data factory with Meltano, dbt, and Podman. 807+ tables and 58M+ rows ingested with an 89% reduction in processing time.

06.2025—08.2025

Data Analyst Intern

Decathlon

Consolidated 14GB from 5 internal systems; the resulting self-serve dashboard was adopted by 9 stakeholders.

07.2024

Data Engineer Intern

AiLand

Worked with large-scale social media data and fine-tuned NLP models for localized social listening.

STACK / WORKING SET

INGESTION Meltano · JDBC · KafkaTRANSFORM dbt · Spark · Python · SQLSERVE FastAPI · Power BI · StreamlitRUN Podman · Docker Compose · Airflow
04writing.feed

READING LOG

Notes from
the pipeline.

Published articles will appear here as a chronological feed once titles, platforms, dates, and live URLs are available.

Idempotency: Why It Matters and How to Achieve It

the most forgotten DE concept in the industry, it is the persisting temporary fix for data specialist.

05contact.endpoint

NEXT CONNECTION

Let's build
the next layer.