m
mfreyeso

Mario R. Ojeda

@mfreyeso

AI Data Engineer

Colombie
Espagnol, Anglais
Certaines informations sont présentées en anglais.
À propos de moi
I am an AI Data Engineer with a Master’s in Artificial Intelligence and over 10 years of experience designing scalable systems. I specialize in data pipeline design, ETL/ELT processes, and workflow orchestration using Python, PostgreSQL, and Apache Airflow across GCP and AWS. I focus on translating complex data challenges into high-value production solutions at the intersection of data engineering, machine learning, and natural language processing.... Plus d’infos

Compétences

m
mfreyeso
Mario R. Ojeda
hors ligne • 
Temps de réponse moyen de 9 heures

Voir mes services

Conseil en technologie de l'IA
I will provide expert ai consulting and data engineering solutions

Expérience professionnelle

Self_employed / Freelance

Senior Data Engineer at Graphite

Self employed / Freelance • Temps plein

Jul 2022 - Feb 2026 • 3 yrs 7 mos

Architected and scaled a multi-source cloud data infrastructure handling 100 TB of data quarterly across 3 distinct database engines, powering the datasets and backend infrastructure behind 90% of all core application features for both internal and external users. Automated 10+ high-throughput ETL/ELT data pipelines using Python and Apache Airflow to ingest and transform Google Search Console (GSC) data, processing over 50M rows daily for dozens of enterprise clients. Optimized data processing workflows to run 3x faster while cutting cloud infrastructure costs 10x by re-engineering pipeline logic to bypass unnecessary executor scaling within Airflow. Engineered scalable ingestion and indexing structures for analytical vector stores, managing ~300M rows with optimized embeddings to achieve sub-400ms synchronous search latencies and highly efficient 300s asynchronous batch lookups using Pinecone, PostgreSQL, and pgvector. Orchestrated large-scale batch ML inference workflows executing 300M predictions quarterly across 4 dedicated production pipelines; achieved a 20x cost reduction compared to traditional inference patterns by leveraging Amazon SageMaker Batch Transform and optimized compute topologies. Implemented automated data quality checks and anomaly detection within CI/CD pipelines, securing data contract compliance and reliability for the downstream tables that power 90% of user-facing application features.

Endava

Senior Software Developer

Endava • Temps plein

Aug 2021 - Jun 2022 • 10 mos

Developed software features as a contractor for Lululemon’s engineering team, contributing to a high-profile production Proof of Concept (POC) deployed across flagship retail stores in Canada. Translated complex business requirements into scalable technical designs, maintaining zero-friction agile release cycles and seamless orchestration aligned with corporate SDLC guidelines. Built serverless backend architectures using Python, AWS Lambda, RDS, and Cognito, supporting iOS mobile applications with optimized APIs designed to handle peak traffic spikes of over 500 concurrent requests per second. Maintained and optimized CI/CD workflows and Infrastructure as Code (IaC) using CloudFormation and Terraform, reducing multi-environment infrastructure deployment times from hours to under 15 minutes.

Self_Employed

Data Engineer at Heinsohn

Self Employed • Temps plein

May 2019 - Jul 2021 • 2 yrs 2 mos

Contributed as a contractor to BrightInsight’s engineering team, designing batch and streaming architectures for two Software as a Medical Device (SaMD) products deployed in initial phases across two major hospital networks. Built high-availability streaming and batch pipelines using Apache Beam, GCP Dataflow, and Apache Airflow to process over 5 million daily telemetric events and device log loops. Implemented asynchronous polling mechanisms in Python to ingest and harmonize EHR data from third-party systems like Epic via Redox, transformationally mapping complex HL7 and FHIR formats into analytical datastores. Designed scalable data schemas using GCP BigQuery and Datastore, establishing architectural patterns that ensured strict HIPAA compliance and zero-leakage security boundaries for over 50,000 active patient profiles. Containerized and deployed a distributed ecosystem of over 10 microservices on Google Kubernetes Engine (GKE), seamlessly orchestrating data pipelines with core identity, access management, and authorization services. Developed the integration layer required to feed clean, pre-processed healthcare datasets into upstream machine learning models within the existing production platform.