Emmanuel Opoku

Data Analyst · pipelines, SQL, and decisions from data

Case studies that move from raw sources to analytics-ready insight — lakehouse engineering, supply-chain BI, and healthcare prediction.

Who I am

Background, education, and what I’m looking for.

Results-driven data analyst with experience across analytics, business intelligence, and process optimization. I turn messy multi-source data into models and narratives decision-makers can use.

Currently completing a Master’s in Data Analytics at the Berlin School of Business and Innovation, with a prior foundation in Chemistry from KNUST and hands-on work spanning healthcare data and operational reporting.

How I work

The ways I ship insight — not a tool dump.

Data engineering

Medallion pipelines, Spark/Delta, notebook orchestration from raw ingest to analytics-ready tables.

SQL & modeling

PostgreSQL exploration, normalization, quality tests, and star-schema design for reporting.

Business intelligence

Interactive Tableau dashboards that surface KPIs, trends, and recommendations for operators.

Applied machine learning

Feature engineering, model comparison, and evaluation with recall-aware choices for real risk.

Case studies

Three projects. Each has a live surface matched to what the work actually is.

01 · Data engineering

Bike Data Lakehouse

Raw CRM & ERP sources → Bronze / Silver / Gold on Databricks, ending in a star schema ready for BI.

  • Databricks
  • Apache Spark
  • Delta Lake
  • Medallion
  • Star schema

Bronze — source of truth

Raw ingested CRM and ERP datasets land here with no business transformations. Preserves lineage before cleansing.

  • Ingest from multi-source files
  • Store as Delta tables
  • Orchestrated via notebook run

Bronze → Silver → Gold

02 · Analytics & BI

DataCo Supply Chain

PostgreSQL analysis of ~180k orders and interactive Tableau dashboards — sales decline, late deliveries, and retention gaps made visible for strategy.

  • PostgreSQL
  • Excel
  • Tableau
  • EDA
  • Dashboard design

03 · Machine learning

Diabetes risk from basic health data

Classifiers trained on age, gender, sugar level, weight, height, and BMI — with SVM preferred for high recall on positive cases.

  • Python
  • scikit-learn
  • SVM · Decision Tree · KNN
  • Feature engineering

Interactive demo of the project interface. Scoring here is a transparent rule layer over the same features used in the notebook (BMI + sugar thresholds) — not the serialized SVM. See the write-up for model metrics and selection rationale.

Result

Enter values to estimate

Uses age, gender, sugar, weight, height → BMI, matching the project feature set.

Let’s talk

Open to Data Analyst and Analytics Engineer opportunities.