Data Science
Transforming complex datasets into predictive insights, intelligent models, and measurable business value.
I design and develop reliable Data and AI solutions: data architecture and processing, model training, API development, and production deployment.
I combine data science, data engineering, and AI engineering to design systems that move seamlessly from raw data to reliable intelligence.
Transforming complex datasets into predictive insights, intelligent models, and measurable business value.
Designing reliable data platforms and scalable pipelines that turn raw data into production-ready assets.
Building production-grade AI systems by connecting models, data, APIs, and cloud infrastructure.
Engineering cloud-native architectures for analytics, observability, automation, and intelligent applications.
Data Science/Data Engineering/Artificial Intelligence
A selection of Data Engineering, Machine Learning, and AI Engineering projects, ranging from real-time data platforms and RAG systems to MLOps, cloud architectures, and intelligent solutions powered by simulation and optimization.
Designed and developed a distributed Data Engineering platform for real-time data ingestion and processing using Apache Airflow, Kafka and Spark Structured Streaming. Implemented Avro data contracts with Schema Registry, data validation and quarantine, streaming analytics and Parquet storage on MinIO/S3. Applied production-oriented engineering with a multi-broker Kafka cluster, distributed Spark cluster, Prometheus/Grafana observability, Kafka authentication and ACLs, automated testing, Docker and CI/CD pipelines with GitHub Actions.
Designed and developed an end-to-end MLOps platform for training, evaluating, serving and monitoring a breast cancer classification model using scikit-learn. Built a reproducible ML pipeline covering data validation and preprocessing, model evaluation, experiment tracking and model management with MLflow. Developed a FastAPI inference service exposing predictions and probabilities, with Prometheus and Grafana for application and model-serving observability. Applied production-oriented engineering practices with Docker, automated unit and integration testing, and a CI/CD pipeline using GitHub Actions.
Designed and developed a multilingual RAG system (FR/EN) using Python, FastAPI, Pydantic, PostgreSQL/pgvector, embeddings, vector retrieval and LLM APIs (Gemini/OpenAI), with streaming responses and structured source citations. Built an end-to-end RAG pipeline and evaluation framework covering retrieval quality (Hit@K, Recall@K, MRR), groundedness, citation correctness, answer completeness and relevance, uncertainty handling and FR/EN multilingual consistency. Applied production-oriented AI/Software Engineering practices including a modular, provider-agnostic architecture, PostgreSQL/pgvector, Docker, Alembic, pytest (575+ tests), Ruff and regression thresholds for automated system quality monitoring.
Designed a cloud-native Data Lakehouse architecture on AWS to centralize, transform, catalog and analyze data from multiple sources. Data is stored in Amazon S3, transformed through ETL pipelines with AWS Glue and organized in a data catalog to improve data discovery and governance. Amazon Athena provides serverless analytics directly on the Data Lake, while Amazon Redshift delivers a dedicated analytical layer for Data Warehouse workloads. The infrastructure is defined and automated with Terraform to provide reproducible deployments and a scalable cloud data architecture.
Designed a digital twin platform for monitoring, simulating and optimizing intelligent energy systems. The architecture combines real-time data acquisition, Machine Learning models and simulation mechanisms to reproduce the behavior of physical assets, analyze their operating state and anticipate system evolution. A FastAPI service provides access to data and models, MongoDB manages operational data, and Three.js enables interactive visualization of equipment and system states. The platform is containerized with Docker to provide a modular and reproducible architecture.
Designed an intelligent engine for forecasting, simulating and optimizing renewable energy production. The solution combines Machine Learning models with TensorFlow, data analysis with Pandas and optimization algorithms to leverage historical and operational data, forecast energy production and identify more efficient operating strategies. A Digital Twin approach enables different scenarios to be simulated before implementation, while AWS services provide the infrastructure required for data storage and processing.
From data pipelines and machine learning systems to AI applications and intelligent digital twins.
A professional journey across data science, software engineering, machine learning, cloud architecture, and MLOps — combining analytical thinking with production-oriented engineering.
Designed and developed end-to-end data and software solutions covering architecture, data engineering, machine learning, MLOps, APIs, cloud infrastructure, and deployment.
Built a conversational RAG assistant with document ingestion, embeddings, Redis Vector Search, FastAPI, LLM tool calling, Pydantic validation, and real-time SSE streaming.
Developed Qardyl, a full-stack AWS application using React, TypeScript, Python, API Gateway, DynamoDB, S3, CloudFront, Cognito, IAM, SES, and GitHub Actions.
Developed Data Science and Machine Learning solutions with a focus on statistical analysis, predictive modeling, and industrial time-series data.
Developed Machine Learning models for Fabemi to identify factors influencing vibrating press performance and predict recipe preparation times.
Contributed to SmartForest through technology selection, data platform development, and monitoring and visualization interfaces.
Developed a SmartForest proof of concept focused on anomaly detection and predictive analysis for industrial systems.
Applied Machine Learning and statistical approaches to identify abnormal operating patterns in industrial data.
Explored predictive maintenance approaches to support early detection of potential equipment failures.
Developed statistical and predictive models for events unrelated to communication and purchasing behavior.
Applied data analysis and statistical modeling techniques to identify patterns in historical datasets.
Worked on forecasting approaches to improve purchasing prediction and inventory management.
Performed statistical analysis and forecasting of births, deaths, and net migration using demographic datasets.
Contributed to demographic projection studies and the analysis of population trends.
Worked on urban air-quality analysis and statistical studies supporting evidence-based decision making.
An academic foundation in statistics and data science, complemented by continuous learning across software engineering, big data, artificial intelligence, and cloud technologies.
Advanced training in data science, statistical modeling, machine learning, data analysis, and quantitative methods.
Training in statistics, probability, mathematical modeling, data analysis, and statistical computing.
Software Development
HTML5, CSS3, JavaScript, React, Node.js, Express, MongoDB, REST APIs and full-stack web development.
Colt Steele
Python Development
Python programming, object-oriented programming, automation, APIs, web development and software development.
Andrei Neagoie
Docker & Kubernetes
Docker, Docker Compose, Multi-Container Projects, Deployment and Kubernetes Fundamentals.
Maximilian Schwarzmüller
Knowledge badges covering core AWS concepts, architecture, core networking, serverless and data migration.
A multidisciplinary stack combining data science, machine learning, AI engineering, data engineering, cloud and software development.
I combine technical depth with an end-to-end engineering approach: transforming raw data into reliable models, intelligent applications, scalable data platforms and production-ready AI systems.
Designing scalable architectures for data, machine learning and AI applications from experimentation to production.
Building data workflows and analytical solutions that transform complex datasets into actionable insights.
Developing predictive models, time-series solutions, anomaly detection and optimization approaches.
Developing LLM, RAG and generative AI systems with retrieval, tool calling and production APIs.
Designing batch and streaming pipelines for reliable ingestion, transformation, processing and monitoring.
Automating model workflows, deployment, monitoring and reproducible machine learning environments.
Designing and deploying scalable cloud solutions using AWS, containers, serverless services and managed infrastructure.
Building maintainable backend services, APIs and full-stack applications around data and AI use cases.
From data foundations to intelligent production systems
Have a project, research idea, data challenge, or AI system in mind? Let's discuss how data, engineering, and intelligent technologies can turn it into a reliable solution.