Principal Software Engineer Solution Architect · Technical Leader

Ibrahim Elmansoury

Systems that endure. Ideas that move.

Ibrahim architects distributed platforms, intelligent products, and engineering teams that turn complicated business problems into dependable systems.

10+ years building software100M+ daily ad impressions20K+ connected players3K+ retail locations

Based in Egypt · Working worldwide · Open to relocation

Understand

Business goals, users, constraints, risks

Selected impact

Systems measured in outcomes.

100M+
daily ad impressionsProgrammatic media delivery at retail scale
99.9%
programmatic delivery successReliable campaign delivery across a distributed retail network
20K+
connected media playersOperating across 3,000+ retail locations
3K+
retail locationsA distributed physical media network
98%+
playback reliabilityReliable delivery under variable connectivity
<100 ms
synchronization driftCoordinated multi-player audio and video
~90%
faster report generationBackground processing and notification workflows
70%
lower query latencyAfter a zero-downtime analytics migration
40+
repositories standardizedUnified tooling, libraries, CI/CD, and delivery practices
1,387
problems solvedDeliberate LeetCode practice begun during the COVID lockdown period
Professional profileIE–2014–NOW

Classification · Principal-level engineering

The case for an engineer who works across the whole system.

Ibrahim moves between architecture, implementation, delivery, and leadership without treating them as separate jobs.

Outcomes over theatre.

  1. 01Architects distributed and cloud-native systems
  2. 02Works deeply in critical-path code
  3. 03Leads cross-functional international teams
  4. 04Turns ambiguity into reliable products
  5. 05Improves developer experience and delivery systems
  6. 06Builds across AI, data, backend, frontend, and platforms
  7. 07Mentors engineers and raises technical standards
Reviewed for accuracyTECHNICAL AUTHORITYOpen file ↓
01System case studies

What was happening, why it was difficult, how the solution worked, what changed, and the part I personally owned.

Project 01

Enterprise advertising platform

What was happening

Deliver programmatic campaigns reliably across a large, distributed physical network without sacrificing product velocity or operational control.

Why it was difficult

Campaign decisions had to reach thousands of physical devices, tolerate unreliable connectivity, and remain observable when delivery failed.

What I designed

I led the platform architecture and long-term technical roadmap, designing cloud-native services around scalable delivery, fault tolerance, observability, and clear service boundaries.

What changed

The platform serves 100M+ daily impressions across 20,000+ players in 3,000+ retail locations with 99.9% programmatic delivery success and 98%+ playback reliability.

After · coordinated delivery
Campaign controltargeting · policy · schedule
Edge deliverycache · retry · observe
Player 01Player 02Player 03
How the solution worked
  • Delivery guarantees over optimistic sends
  • Operational visibility across a physical network
  • Scale without coupling campaign logic to devices
The deliberate compromise

More delivery state and operational discipline in exchange for reliability at retail scale.

My role

I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.

Project 02

In-store video synchronization

What was happening

Play the same video across multiple screens inside a store while networks are unreliable and large media assets are expensive to move repeatedly.

Why it was difficult

Every screen had its own clock and buffering behavior. The design had to correct drift continuously, survive a lost connection, and avoid downloading the same large video repeatedly.

What I designed

I designed scheduled playback epochs, WebSocket timing signals, local media prefetching, player-clock telemetry, continuous drift correction, and offline playback from the store edge.

What changed

Video playback across in-store screens stays below 100 ms drift, while adaptive delivery and edge caching reduced bandwidth consumption by approximately 60%.

In-store video synchronization
Control planeSchedule + playback epochcampaign rules · store groups · media manifest
WebSocket timing signal
Store edgeprefetched media · local cache
Screen 01clock +0 ms
Screen 02clock +42 ms
Screen 03clock −31 ms
telemetrymeasure driftcorrect playback
<100 msplayback drift across in-store screensOffline? Continue from the local schedule, then reconcile.
How the solution worked
  • WebSocket timing signals coordinate a shared playback epoch
  • Edge prefetching trades local storage for lower bandwidth
  • Clock telemetry drives correction without visibly restarting video
The deliberate compromise

Local device complexity in exchange for continuity, synchronization, and lower bandwidth.

My role

I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.

Project 03

Analytics workflow modernization

What was happening

Synchronous report generation blocked user workflows and became too slow as analytical volume and reporting complexity grew.

Why it was difficult

The workflow had to execute expensive analytical queries and generate large files without holding an HTTP request open, duplicating jobs, losing progress, or hiding failures.

What I designed

I designed a durable asynchronous workflow that validates requests, persists job state, queues work, processes partitions in parallel, queries analytical stores, streams generated artifacts into object storage, and notifies users when results are ready. Retries, dead-letter handling, and job-level observability make failures recoverable.

What changed

Report-generation time fell by approximately 90%, with a clearer and more resilient user workflow.

Asynchronous analytics workflow
01 · APIValidate requestAccepted immediately
  1. 02Persist jobstatus · owner · filters
  2. 03Durable queuepriority · retry policy
  3. 04Worker poolparallel partitions
  4. 05Query layerwarehouse · cache
  5. 06Build artifactstream · compress
  6. 07Object storagesigned result URL
  7. 08Notify userin-app · email
Retries exponential backoffDead letter failed jobs isolatedObservability duration · state · errors
Before · blocked request~90% fasterAfter · durable background workflow
How the solution worked
  • Persist job state before queueing expensive work
  • Parallel workers generate and stream artifacts without holding the request
  • Retries, dead letters, and notifications make failure recoverable
The deliberate compromise

More workflow and job-state machinery in exchange for a faster, more resilient reporting experience.

My role

I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.

Project 04

Data platform re-architecture

What was happening

Operational and analytical workloads competed for resources, producing slow queries while the business required uninterrupted service.

Why it was difficult

The migration had to separate analytical work from production traffic without downtime or a risky one-time cutover.

What I designed

I re-architected the analytical path and led a zero-downtime migration using Airflow, PySpark, BigQuery, TimescaleDB, and materialized views.

What changed

Analytical query latency fell by approximately 70% without interrupting production.

Before
Operational DBCoupled analyticsSlow queries
Contention · long feedback loops
After
SourcesAirflow + PySparkAnalytical stores
BigQuery · TimescaleDB · materialized views
How the solution worked
  • Separate operational and analytical workloads
  • Migrate without a business-visible cutover
  • Design orchestration for repeatability and recovery
The deliberate compromise

Migration and pipeline complexity in exchange for faster analytics and safer evolution.

My role

I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.

02The operating range

I move from business ambiguity into implementation, production, and team leverage—without floating above the work.

What I do at this layer

Implementation

  • APIs and distributed services
  • Data models and pipelines
  • Critical code and debugging
03Experience · career timeline

Consultant, technical lead, senior engineer, principal engineer, solution architect, and Head of Engineering—each role expands the scope while preserving hands-on ownership.

01
Heru Loop

Head of Engineering · Solution Architect

I lead the technical direction of an IT services and product-engineering company that helps organisations move from an ambiguous idea to reliable production software.

Jul 2026 — PresentEgypt · Remote

Shape product and technology strategy across digital platforms, AI-enabled systems, workflow automation, and scalable custom software.

Lead technical discovery with clients, translating business constraints into architecture, phased delivery plans, and explicit engineering trade-offs.

Remain hands-on in system design, implementation decisions, code review, infrastructure, and the critical paths that carry the most delivery risk.

Read details +
Technical strategyHands-on architectureProduct engineeringClient delivery
  • Shape product and technology strategy across digital platforms, AI-enabled systems, workflow automation, and scalable custom software.
  • Lead technical discovery with clients, translating business constraints into architecture, phased delivery plans, and explicit engineering trade-offs.
  • Remain hands-on in system design, implementation decisions, code review, infrastructure, and the critical paths that carry the most delivery risk.
  • Build delivery practices around reusable foundations, clear ownership, observability, security, and maintainable operations.
  • Align engineering work with real workflows and measurable client outcomes instead of treating technology as an isolated deliverable.

From business ambiguity to production systems

Problem

Clients often arrive with a valuable business goal but no shared model of the users, workflows, data, risks, or technical path required to deliver it.

What I changed

I lead discovery, solution architecture, and engineering execution as one connected process—keeping business decisions, system constraints, and implementation feedback visible to the team.

Outcome

Projects move forward with clearer scope, deliberate trade-offs, and a technical foundation designed to survive beyond the first release.

  • Product engineering
  • AI-enabled systems
  • Automation
  • Cloud platforms
02
Intouch.com

Principal Software Engineer · Solution Architect

I own technical direction for high-scale media, analytics, and AI-enabled engineering systems while leading a cross-functional team of 10+ engineers.

Jul 2023 — PresentEl Gouna, Egypt · Ireland

Own the technical vision and long-term evolution of a programmatic media platform serving 100M+ daily impressions.

Lead and mentor 10+ engineers through architecture reviews, design guidance, incident analysis, and continuous technical feedback.

Mentor engineers across cross-functional squads through regular 1:1s, performance reviews, promotion cycles, growth plans, and clear technical feedback.

Read details +
Hands-on ICSystem architectureTechnical leadershipDelivery management
  • Own the technical vision and long-term evolution of a programmatic media platform serving 100M+ daily impressions.
  • Lead and mentor 10+ engineers through architecture reviews, design guidance, incident analysis, and continuous technical feedback.
  • Mentor engineers across cross-functional squads through regular 1:1s, performance reviews, promotion cycles, growth plans, and clear technical feedback.
  • Partner with product managers, enterprise customers, and business stakeholders during discovery, solution design, and delivery planning.
  • Write RFCs, ADRs, migration plans, and technical specifications that turn cross-team decisions into executable work.
  • Balance reliability, scale, operational simplicity, product velocity, delivery dependencies, and business priorities across concurrent initiatives.
  • Stay hands-on in distributed services, data pipelines, real-time media, AI workflows, platform tooling, and production investigations.

Enterprise advertising platform

Problem

Deliver programmatic campaigns reliably across a large, distributed physical network without sacrificing product velocity or operational control.

What I changed

I led the platform architecture and long-term technical roadmap, designing cloud-native services around scalable delivery, fault tolerance, observability, and clear service boundaries.

Outcome

The platform serves 100M+ daily impressions across 20,000+ players in 3,000+ retail locations with 99.9% programmatic delivery success and 98%+ playback reliability.

  • Distributed services
  • Event-driven architecture
  • Caching
  • AWS · GCP · Azure

In-store video synchronization

Problem

Play the same video across multiple screens inside a store while networks are unreliable and large media assets are expensive to move repeatedly.

What I changed

I designed scheduled playback epochs, WebSocket timing signals, local media prefetching, player-clock telemetry, continuous drift correction, and offline playback from the store edge.

Outcome

Video playback across in-store screens stays below 100 ms drift, while adaptive delivery and edge caching reduced bandwidth consumption by approximately 60%.

  • Video playback
  • WebSockets
  • Clock synchronization
  • Offline continuity
  • Edge caching

Analytics workflow modernization

Problem

Synchronous report generation blocked user workflows and became too slow as analytical volume and reporting complexity grew.

What I changed

I designed a durable asynchronous workflow that validates requests, persists job state, queues work, processes partitions in parallel, queries analytical stores, streams generated artifacts into object storage, and notifies users when results are ready. Retries, dead-letter handling, and job-level observability make failures recoverable.

Outcome

Report-generation time fell by approximately 90%, with a clearer and more resilient user workflow.

  • Background processing
  • Job orchestration
  • Notifications
  • Analytics

Data platform re-architecture

Problem

Operational and analytical workloads competed for resources, producing slow queries while the business required uninterrupted service.

What I changed

I re-architected the analytical path and led a zero-downtime migration using Airflow, PySpark, BigQuery, TimescaleDB, and materialized views.

Outcome

Analytical query latency fell by approximately 70% without interrupting production.

  • Airflow
  • PySpark
  • BigQuery
  • TimescaleDB
  • Materialized views
03
Intouch.com

Senior Software Engineer

I built high-throughput backend and data systems, improved production reliability, and helped modernize the platform around reusable services and engineering standards.

May 2022 — Jun 2023Egypt

Built a .NET batch-processing system that reduced a 1M+ record import from three hours to 25 minutes.

Developed backend platforms for campaign management, media delivery, analytics, reporting, and enterprise workflows.

Built Airflow-orchestrated ETL pipelines across cloud analytics platforms.

Read details +
Hands-on ICBackend systemsData engineering
  • Built a .NET batch-processing system that reduced a 1M+ record import from three hours to 25 minutes.
  • Developed backend platforms for campaign management, media delivery, analytics, reporting, and enterprise workflows.
  • Built Airflow-orchestrated ETL pipelines across cloud analytics platforms.
  • Removed memory leaks and improved database, cache, and distributed-service performance in production.
  • Created reusable backend frameworks for advanced filtering, pagination, search, sorting, and external API integration.

High-throughput batch processing

Problem

Large record-processing workloads were slow and difficult to operate predictably.

What I changed

I built a resource-efficient .NET batch-processing system designed for million-record imports and predictable resource use.

Outcome

A 1M+ record import fell from three hours to 25 minutes.

  • .NET
  • Batch processing
  • Performance optimization
  • Concurrent processing
04
HeartAttack

Tech Lead · Individual Contributor

I led web, mobile, and backend delivery while continuing to design and implement the systems behind customer products.

Apr 2020 — Jan 2022Remote

Owned technical delivery from solution design and planning through production deployment and operational support.

Coordinated frontend, backend, and mobile engineers while managing dependencies, release risks, and incremental delivery.

Established reusable components, shared libraries, testing practices, documentation, Git workflows, and CI/CD standards.

Read details +
Hands-on ICTechnical directionCross-functional delivery
  • Owned technical delivery from solution design and planning through production deployment and operational support.
  • Coordinated frontend, backend, and mobile engineers while managing dependencies, release risks, and incremental delivery.
  • Established reusable components, shared libraries, testing practices, documentation, Git workflows, and CI/CD standards.
  • Built React and React Native products, reusable design-system components, and data-intensive D3.js dashboards.
  • Designed payment, subscription, and commercial workflow services.
  • Optimized GraphQL and MongoDB access patterns for a platform processing 10,000+ orders per day.
  • Used Go for backend services and concurrent workloads where predictable performance and straightforward deployment mattered.

Commercial product platforms

Problem

Product teams needed reliable payment, subscription, analytics, and order workflows that could evolve quickly.

What I changed

I aligned product, design, and engineering around an incremental architecture; built backend services and React products; and improved GraphQL and MongoDB performance.

Outcome

The platform supported 10,000+ orders per day with stronger performance and operational reliability.

  • Go
  • React · React Native
  • GraphQL
  • MongoDB
  • Payments
05
Independent

Software Engineering Consultant

I partnered with startups and growing businesses from discovery through production across commerce, logistics, mobility, and enterprise operations.

Sep 2016 — Apr 2020Egypt · Remote

Led customer discovery, architecture workshops, technical planning, implementation, deployment, and production support.

Modernized legacy platforms through cloud adoption, API integrations, clearer service boundaries, and performance optimization.

Improved production systems using database tuning, asynchronous processing, caching, validation, monitoring, and error handling.

Read details +
DiscoveryArchitectureImplementationProduction ownership
  • Led customer discovery, architecture workshops, technical planning, implementation, deployment, and production support.
  • Modernized legacy platforms through cloud adoption, API integrations, clearer service boundaries, and performance optimization.
  • Improved production systems using database tuning, asynchronous processing, caching, validation, monitoring, and error handling.
  • Connected product software to payments, subscriptions, ERP systems, logistics workflows, live tracking, and operational reporting.
  • Worked directly with founders, executives, and operational stakeholders to turn manual processes into maintainable software.
  • Used Go alongside Python, Java, PHP, and JavaScript to build efficient services and integrations suited to each client system.

Systems shaped around the business

Problem

Teams had operational bottlenecks, legacy platforms, and product ideas without a dependable technical path.

What I changed

I designed and delivered platforms across commerce, vehicle rental, agricultural logistics, ride-hailing, shipping, and enterprise operations—combining Go and other backend technologies with APIs, integrations, payments, real-time workflows, and cloud modernization.

Outcome

Manual workflows became software products with clearer service boundaries, faster data access, and more reliable operations.

  • Go
  • FastAPI
  • Laravel
  • Spring Boot
  • Real-time products
04Ventures · classified advertisements

Education, developer tooling, and AI-era software engineering—each has a distinct problem, audience, and operating model.

Service AeventQueue Service BFailure → retry or compensate?
Learning platformIn development

SystemCraft

An interactive learning platform that makes distributed systems and software architecture approachable through scenarios, trade-offs, visual explanations, and hands-on exploration.

Why it exists

Software engineers often learn patterns by name before they learn when those patterns help, fail, or create new trade-offs.

  • Distributed systems
  • Consistency
  • Time and ordering
  • Sagas
  • Reliability
  • Architecture trade-offs
Currently being built · details available on request
reporeporepo
Shikamaru
DockerenvAzure
Developer toolingOpen source · Published on npm

Shikamaru

Developer tooling that turns repetitive engineering workflows into dependable automation across repositories, local infrastructure, and Azure DevOps.

Why it exists

Engineers lose attention switching between repositories, environment setup, service dependencies, and repetitive platform operations.

  • Multi-repo automation
  • Docker environments
  • Azure DevOps synchronization
Read the documentation
DiscoverReasonBuildObserveLearnHuman judgment
AI-era software engineering companyStudio initiative

Heru Loop

An AI-native software studio building intelligent workflows, agentic systems, living products, and the reliable platforms behind them.

Why it exists

Useful AI products need more than a model call: they need real workflow integration, strong engineering, human-centered automation, and feedback loops.

  • Systems, not isolated screens
  • Workflow-connected AI
  • Discovery to production
  • Continuous feedback
Currently being built · details available on request
Open source · Technical education

Engineering leverage that compounds.

I publish tools and learning material when repeated friction reveals a problem worth solving for other engineers.

1,387algorithmic problems solved through deliberate practiceView LeetCode ↗
Public · Available on npm

Shikamaru CLI

A command-line toolkit for automating multi-repository environments, Docker-based infrastructure setup, and Azure DevOps synchronization.

I built it to reduce the operational and cognitive overhead that accumulates when developers work across many services and environments.
Documentationnpm package
Internal foundations

Reusable engineering systems

Shared libraries, backend frameworks, delivery pipelines, observability patterns, and development standards.

Built to reduce duplicated decisions and focus teams on product-specific problems.
Lockdown practice notes

How I spent the quiet part of COVID.

I worked systematically through LeetCode to strengthen how I decompose unfamiliar problems, test assumptions, reason about trade-offs, and translate an approach into efficient code.

291
Easy
797
Medium
299
Hard
5.8K
Submissions

1,387 problems solved. The habit outlasted lockdown: stay patient with ambiguity and keep improving the model.

Open my LeetCode profile ↗
LeetCode profile showing 1,387 solved problems: 291 easy, 797 medium, and 299 hard, with 5.8 thousand submissions
Profile snapshot · July 2026
05 · Continuous evolution

Learning feeds the systems I build.

Formal study, production incidents, technical reading, teaching, and deliberate practice become better engineering judgment.

Formal education

MSc in Artificial Intelligence

Woolf University · LLMs, machine-learning systems, scalable AI, and applied AI engineering.

Current themes
  • Agentic AI systems
  • Distributed systems
  • Data platforms
  • Reliability engineering
  • Architecture and technical leadership

SystemCraft turns these themes into public technical education.

Operating agreementTerms & conditions apply in production

Principles I use when the system gets difficult.

  1. §1Systems, not screens
  2. §2Understand the problem before choosing the technology
  3. §3Reliability is a product feature
  4. §4Architecture must survive contact with production
  5. §5Make complexity visible
  6. §6Automate repeated engineering work
  7. §7Use AI to amplify judgment, not replace it
  8. §8Own the outcome, not only the ticket

Agreed — Ibrahim

Technical capabilities · grouped by problem

Open the relevant file.

01Distributed systems+

Systems that remain dependable as traffic, locations, teams, and failure modes grow.

Tools and methods

Event-driven architecture · Messaging · Coordination · Caching · Retries · Delivery guarantees

Where I used it

Applied across programmatic media delivery, synchronized playback, offline continuity, and fault-tolerant cloud services.

02Backend and APIs+

Service boundaries, APIs, and processing workflows designed for clear ownership and predictable production behavior.

Tools and methods

TypeScript, Node.js, NestJS, Go, Python, Java, .NET · REST, GraphQL, gRPC · RabbitMQ, Kafka

Where I used it

Built high-throughput processing, payment and subscription services, enterprise integrations, and reusable backend foundations.

03AI and agentic systems+

Reliable AI-enabled workflows connected to real knowledge, users, and operational constraints—not isolated model demos.

Tools and methods

LLM integration · RAG · Embeddings and retrieval · Agent memory and tools · Evaluation · Human-in-the-loop design

Where I used it

Applied to text-to-speech advertising, campaign recommendations, enterprise knowledge retrieval, and engineering documentation.

04Data platforms+

Data movement and analytical models designed around query patterns, freshness, cost, and production operability.

Tools and methods

PostgreSQL, MongoDB, Redis, Elasticsearch · BigQuery, Airflow, Spark, PySpark · Batch and streaming pipelines

Where I used it

Led a zero-downtime analytics modernization that reduced query latency by 70%, alongside enterprise ETL and reporting platforms.

05Frontend and product engineering+

Product workflows built end to end, with enough frontend depth to connect system behavior to the user experience.

Tools and methods

Angular, React, Next.js, React Native · WebSockets · Progressive web applications · Accessibility and responsive design

Where I used it

Delivered web and mobile products, data-intensive dashboards, reusable design-system components, and real-time customer workflows.

06Cloud and DevOps+

Delivery systems and observability treated as part of the product—not cleanup after implementation.

Tools and methods

AWS, GCP, Azure · Docker, Kubernetes, Terraform · CI/CD · OpenTelemetry, Prometheus, Loki · Incident response

Where I used it

Owned serverless media pipelines, multi-cloud analytics, standardized delivery across 40+ repositories, and long-term reliability improvements.

07Reliability and observability+

Production behavior made visible enough to diagnose failures, learn from incidents, and improve the system deliberately.

Tools and methods

OpenTelemetry · Prometheus · Loki · Structured logging · Alerting · Incident analysis · Performance profiling

Where I used it

Improved playback continuity, removed memory leaks, diagnosed distributed-service bottlenecks, and built observable delivery workflows.

08Architecture and technical strategy+

Business goals and constraints translated into system boundaries, explicit decisions, migration paths, and executable plans.

Tools and methods

Technical discovery · RFCs · ADRs · Architecture reviews · Migration planning · Roadmaps · Risk management

Where I used it

Owned long-term platform direction, customer solution design, zero-downtime migrations, and cross-team architecture decisions.

09Engineering leadership+

Teams supported with clear context, useful feedback, strong technical standards, and ownership of outcomes.

Tools and methods

Cross-functional squads · Mentoring · 1:1s · Performance reviews · Promotion cycles · Delivery leadership

Where I used it

Led and mentored international teams of 10+ engineers while remaining hands-on in architecture, code, and production.

Portrait of Ibrahim Elmansoury
Available worldwide
Late-night systems counsel · No billable hour required

Got a difficult system? Ibrahim can help make the case.

Ibrahim is open to Principal Engineer, Solution Architect, Staff Platform, and technical-leadership opportunities worldwide—including roles offering visa sponsorship or relocation.

Send an email ↗LinkedIn ↗GitHub ↗View résumé ↗