Principal Software Engineer Solution Architect · Technical Leader

Ibrahim Elmansoury

Systems that endure. Ideas that move.

Ibrahim architects distributed platforms, intelligent products, and engineering teams that turn complicated business problems into dependable systems.

Discuss a problem
10+ years building dependable softwareAI products from discovery to productionRAG systems refined for reliable retrievalDistributed platforms operating at global scaleCross-functional engineering leadership

Based in Egypt · Working worldwide · Open to relocation

The Engineering Desk · Field Notes

How complicated systems become dependable.

By Ibrahim Elmansoury · Principal Engineer

Part 01 of 05
The central thesisEngineering coreStrategy meets execution
ProductPlatformData
01Understand

Turn an unclear business problem into explicit goals, constraints, risks, and measurable outcomes.

Discovery mapNon-functional requirementsRisk registerSuccess metrics
Career-wide impact

Impact across systems, AI products, and teams.

10+ yrs
building production softwareArchitecture, implementation, delivery, and long-term system evolution
AI → prod
intelligent products deliveredFrom workflow discovery and model integration to dependable production systems
RAG
retrieval systems refinedImproving grounding, retrieval quality, evaluation, and operational feedback loops
10+
engineers led and mentoredInternational cross-functional squads, 1:1s, performance reviews, and promotion cycles
100M+
daily reach at retail scale20K+ players · 3K+ locations · 99.9% delivery · <100 ms synchronization
70–90%
data workflows accelerated~70% lower analytical query latency · ~90% faster report generation
Professional profileIE–2014–NOW

Classification · Principal-level engineering

The case for an engineer who works across the whole system.

Ibrahim moves between architecture, implementation, delivery, and leadership without treating them as separate jobs.

Outcomes over theatre.

  1. 01Architects distributed and cloud-native systems
  2. 02Works deeply in critical-path code
  3. 03Leads cross-functional international teams
  4. 04Turns ambiguity into reliable products
  5. 05Improves developer experience and delivery systems
  6. 06Builds across AI, data, backend, frontend, and platforms
  7. 07Mentors engineers and raises technical standards
Reviewed for accuracyTECHNICAL AUTHORITYOpen file ↓
01System case studies

What was happening, why it was difficult, how the solution worked, what changed, and the part I personally owned.

Selected case01 / 4
Project 01

Enterprise advertising platform

Distributed servicesEvent-driven architectureCachingAWS · GCP · Azure
What was happening

Deliver programmatic campaigns reliably across a large, distributed physical network without sacrificing product velocity or operational control.

Why it was difficult

Campaign decisions had to reach thousands of physical devices, tolerate unreliable connectivity, and remain observable when delivery failed.

What I designed

I led the platform architecture and long-term technical roadmap, designing cloud-native services around scalable delivery, fault tolerance, observability, and clear service boundaries.

What changed

The platform serves 100M+ daily impressions across 20,000+ players in 3,000+ retail locations with 99.9% programmatic delivery success and 98%+ playback reliability.

After · coordinated delivery
Campaign controltargeting · policy · schedule
Edge deliverycache · retry · observe
Player 01Player 02Player 03
How the solution worked
  • Delivery guarantees over optimistic sends
  • Operational visibility across a physical network
  • Scale without coupling campaign logic to devices
The deliberate compromise

More delivery state and operational discipline in exchange for reliability at retail scale.

My role

I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.

02Applied AI product use cases

Two product concepts that combine generative AI with professional controls, transparent reasoning, and dependable system design.

Use case 01 · Applied generative AI

Photo direction through conversation.

A professional image workspace where a user can describe intent in natural language, compare revisions, and make precise changes directly on the canvas.

What was happening

Creative teams could generate images quickly, but prompt-only tools made precise direction, repeatable revisions, and brand-safe review difficult.

Why it was difficult

Natural-language intent had to coexist with canvas-level control, protected product details, asynchronous generation, and a reviewable history.

What Ibrahim designed

A conversational workspace that turns requests into structured edits while masks, versioning, approval, and export remain explicit.

What changes

Teams move faster without surrendering the precision, traceability, and human approval required for professional creative work.

Creative session3 revisions
User

Make it feel more premium. Keep the label readable.

AI collaborator

I reduced visual noise, added a controlled rim light, and protected the label as a locked region.

Describe the next change…
Campaign / Product 01
1:1100%Compare
HERUFORM / 01
Locked product

Premium studio

✓ Brand label protected2048 × 2048 · ready to export
How the solution works

The conversation becomes structured edit instructions. The canvas applies them non-destructively, preserves locked regions, and keeps every revision reproducible.

Ibrahim's role

Ibrahim shaped the product model, interaction between conversation and canvas, AI orchestration boundaries, production controls, and human approval path.

The deliberate compromise

Model orchestration, asset storage, safety checks, version history, asynchronous generation, and human approval add machinery—but turn a prompt box into a dependable workflow.

03The operating range

I move from business ambiguity into implementation, production, and team leverage—without floating above the work.

What I do at this layer

Implementation

  • APIs and distributed services
  • Data models and pipelines
  • Critical code and debugging
04Experience · career timeline

Consultant, technical lead, senior engineer, principal engineer, solution architect, and Solutions AI Architect—each role expands the scope while preserving hands-on ownership.

01
Heru Loop

Solutions AI Architect

I lead the technical direction of an IT services and product-engineering company that helps organisations move from an ambiguous idea to reliable production software.

May 2026 — PresentRemote

Shape product and technology strategy across digital platforms, AI-enabled systems, workflow automation, and scalable custom software.

Lead technical discovery with clients, translating business constraints into architecture, phased delivery plans, and explicit engineering trade-offs.

Remain hands-on in system design, implementation decisions, code review, infrastructure, and the critical paths that carry the most delivery risk.

Read details +
Full-timeTechnical strategyHands-on architectureProduct engineeringClient delivery
  • Shape product and technology strategy across digital platforms, AI-enabled systems, workflow automation, and scalable custom software.
  • Lead technical discovery with clients, translating business constraints into architecture, phased delivery plans, and explicit engineering trade-offs.
  • Remain hands-on in system design, implementation decisions, code review, infrastructure, and the critical paths that carry the most delivery risk.
  • Build delivery practices around reusable foundations, clear ownership, observability, security, and maintainable operations.
  • Align engineering work with real workflows and measurable client outcomes instead of treating technology as an isolated deliverable.

From business ambiguity to production systems

Problem

Clients often arrive with a valuable business goal but no shared model of the users, workflows, data, risks, or technical path required to deliver it.

What I changed

I lead discovery, solution architecture, and engineering execution as one connected process—keeping business decisions, system constraints, and implementation feedback visible to the team.

Outcome

Projects move forward with clearer scope, deliberate trade-offs, and a technical foundation designed to survive beyond the first release.

  • Product engineering
  • AI-enabled systems
  • Automation
  • Cloud platforms
02
Intouch.com

Principal Software Engineer · Solution Architect

I own technical direction for high-scale media, analytics, and AI-enabled engineering systems while leading a cross-functional team of 10+ engineers.

Jul 2023 — May 2026Ireland

Own the technical vision and long-term evolution of a programmatic media platform serving 100M+ daily impressions.

Lead and mentor 10+ engineers through architecture reviews, design guidance, incident analysis, and continuous technical feedback.

Mentor engineers across cross-functional squads through regular 1:1s, performance reviews, promotion cycles, growth plans, and clear technical feedback.

Read details +
Hands-on ICSystem architectureTechnical leadershipDelivery management
  • Own the technical vision and long-term evolution of a programmatic media platform serving 100M+ daily impressions.
  • Lead and mentor 10+ engineers through architecture reviews, design guidance, incident analysis, and continuous technical feedback.
  • Mentor engineers across cross-functional squads through regular 1:1s, performance reviews, promotion cycles, growth plans, and clear technical feedback.
  • Partner with product managers, enterprise customers, and business stakeholders during discovery, solution design, and delivery planning.
  • Write RFCs, ADRs, migration plans, and technical specifications that turn cross-team decisions into executable work.
  • Balance reliability, scale, operational simplicity, product velocity, delivery dependencies, and business priorities across concurrent initiatives.
  • Stay hands-on in distributed services, data pipelines, real-time media, AI workflows, platform tooling, and production investigations.

Enterprise advertising platform

Problem

Deliver programmatic campaigns reliably across a large, distributed physical network without sacrificing product velocity or operational control.

What I changed

I led the platform architecture and long-term technical roadmap, designing cloud-native services around scalable delivery, fault tolerance, observability, and clear service boundaries.

Outcome

The platform serves 100M+ daily impressions across 20,000+ players in 3,000+ retail locations with 99.9% programmatic delivery success and 98%+ playback reliability.

  • Distributed services
  • Event-driven architecture
  • Caching
  • AWS · GCP · Azure

In-store video synchronization

Problem

Play the same video across multiple screens inside a store while networks are unreliable and large media assets are expensive to move repeatedly.

What I changed

I designed scheduled playback epochs, WebSocket timing signals, local media prefetching, player-clock telemetry, continuous drift correction, and offline playback from the store edge.

Outcome

Video playback across in-store screens stays below 100 ms drift, while adaptive delivery and edge caching reduced bandwidth consumption by approximately 60%.

  • Video playback
  • WebSockets
  • Clock synchronization
  • Offline continuity
  • Edge caching

Analytics workflow modernization

Problem

Synchronous report generation blocked user workflows and became too slow as analytical volume and reporting complexity grew.

What I changed

I designed a durable asynchronous workflow that validates requests, persists job state, queues work, processes partitions in parallel, queries analytical stores, streams generated artifacts into object storage, and notifies users when results are ready. Retries, dead-letter handling, and job-level observability make failures recoverable.

Outcome

Report-generation time fell by approximately 90%, with a clearer and more resilient user workflow.

  • Background processing
  • Job orchestration
  • Notifications
  • Analytics

Data platform re-architecture

Problem

Operational and analytical workloads competed for resources, producing slow queries while the business required uninterrupted service.

What I changed

I re-architected the analytical path and led a zero-downtime migration using Airflow, PySpark, BigQuery, TimescaleDB, and materialized views.

Outcome

Analytical query latency fell by approximately 70% without interrupting production.

  • Airflow
  • PySpark
  • BigQuery
  • TimescaleDB
  • Materialized views
03
Intouch.com

Senior Software Engineer

I built high-throughput backend and data systems, improved production reliability, and helped modernize the platform around reusable services and engineering standards.

May 2022 — Jun 2023Ireland

Built a .NET batch-processing system that reduced a 1M+ record import from three hours to 25 minutes.

Developed backend platforms for campaign management, media delivery, analytics, reporting, and enterprise workflows.

Built Airflow-orchestrated ETL pipelines across cloud analytics platforms.

Read details +
Hands-on ICBackend systemsData engineering
  • Built a .NET batch-processing system that reduced a 1M+ record import from three hours to 25 minutes.
  • Developed backend platforms for campaign management, media delivery, analytics, reporting, and enterprise workflows.
  • Built Airflow-orchestrated ETL pipelines across cloud analytics platforms.
  • Removed memory leaks and improved database, cache, and distributed-service performance in production.
  • Created reusable backend frameworks for advanced filtering, pagination, search, sorting, and external API integration.

High-throughput batch processing

Problem

Large record-processing workloads were slow and difficult to operate predictably.

What I changed

I built a resource-efficient .NET batch-processing system designed for million-record imports and predictable resource use.

Outcome

A 1M+ record import fell from three hours to 25 minutes.

  • .NET
  • Batch processing
  • Performance optimization
  • Concurrent processing
04
HeartAttack

Tech Lead · Individual Contributor

I led web, mobile, and backend delivery while continuing to design and implement the systems behind customer products.

Apr 2020 — Jan 2022Remote

Owned technical delivery from solution design and planning through production deployment and operational support.

Coordinated frontend, backend, and mobile engineers while managing dependencies, release risks, and incremental delivery.

Established reusable components, shared libraries, testing practices, documentation, Git workflows, and CI/CD standards.

Read details +
Hands-on ICTechnical directionCross-functional delivery
  • Owned technical delivery from solution design and planning through production deployment and operational support.
  • Coordinated frontend, backend, and mobile engineers while managing dependencies, release risks, and incremental delivery.
  • Established reusable components, shared libraries, testing practices, documentation, Git workflows, and CI/CD standards.
  • Built React and React Native products, reusable design-system components, and data-intensive D3.js dashboards.
  • Designed payment, subscription, and commercial workflow services.
  • Optimized GraphQL and MongoDB access patterns for a platform processing 10,000+ orders per day.
  • Used Go for backend services and concurrent workloads where predictable performance and straightforward deployment mattered.

Commercial product platforms

Problem

Product teams needed reliable payment, subscription, analytics, and order workflows that could evolve quickly.

What I changed

I aligned product, design, and engineering around an incremental architecture; built backend services and React products; and improved GraphQL and MongoDB performance.

Outcome

The platform supported 10,000+ orders per day with stronger performance and operational reliability.

  • Go
  • React · React Native
  • GraphQL
  • MongoDB
  • Payments
05
Independent

Software Engineering Consultant

I partnered with startups and growing businesses from discovery through production across commerce, logistics, mobility, and enterprise operations.

Sep 2016 — Apr 2020Egypt · Remote

Led customer discovery, architecture workshops, technical planning, implementation, deployment, and production support.

Modernized legacy platforms through cloud adoption, API integrations, clearer service boundaries, and performance optimization.

Improved production systems using database tuning, asynchronous processing, caching, validation, monitoring, and error handling.

Read details +
DiscoveryArchitectureImplementationProduction ownership
  • Led customer discovery, architecture workshops, technical planning, implementation, deployment, and production support.
  • Modernized legacy platforms through cloud adoption, API integrations, clearer service boundaries, and performance optimization.
  • Improved production systems using database tuning, asynchronous processing, caching, validation, monitoring, and error handling.
  • Connected product software to payments, subscriptions, ERP systems, logistics workflows, live tracking, and operational reporting.
  • Worked directly with founders, executives, and operational stakeholders to turn manual processes into maintainable software.
  • Used Go alongside Python, Java, PHP, and JavaScript to build efficient services and integrations suited to each client system.

Systems shaped around the business

Problem

Teams had operational bottlenecks, legacy platforms, and product ideas without a dependable technical path.

What I changed

I designed and delivered platforms across commerce, vehicle rental, agricultural logistics, ride-hailing, shipping, and enterprise operations—combining Go and other backend technologies with APIs, integrations, payments, real-time workflows, and cloud modernization.

Outcome

Manual workflows became software products with clearer service boundaries, faster data access, and more reliable operations.

  • Go
  • FastAPI
  • Laravel
  • Spring Boot
  • Real-time products
05Ventures · classified advertisements

Education, developer tooling, and AI-era software engineering—each has a distinct problem, audience, and operating model.

Systems Unboxed, SystemCraft visual identity
Service AeventQueue Service B
Open the system. Change one decision. Watch the consequences.
Learning platformBeta · Public preview

SystemCraftBeta

An interactive learning platform that makes distributed systems and software architecture approachable through scenarios, trade-offs, visual explanations, and hands-on exploration.

Why it exists

Software engineers often learn patterns by name before they learn when those patterns help, fail, or create new trade-offs.

  • Distributed systems
  • Consistency
  • Time and ordering
  • Sagas
  • Reliability
  • Architecture trade-offs
Explore the beta
reporeporepo
Shikamaru logoShikamaruOne command · many moving parts
DockerenvAzure
Developer toolingOpen source · Published on npm

Shikamaru

Developer tooling that turns repetitive engineering workflows into dependable automation across repositories, local infrastructure, and Azure DevOps.

Why it exists

Engineers lose attention switching between repositories, environment setup, service dependencies, and repetitive platform operations.

  • Multi-repo automation
  • Docker environments
  • Azure DevOps synchronization
Read the documentation
Human-ledJudgmentsets intent
and approves outcomes
01DiscoverNeed + context
02ReasonModel + decide
03BuildShip the system
04ObserveSignals + outcomes
05LearnImprove the loop

AI accelerates each stage. People remain accountable for the direction.

AI-era software engineering companyStudio initiative

Heru Loop

An AI-native software studio building intelligent workflows, agentic systems, living products, and the reliable platforms behind them.

Why it exists

Useful AI products need more than a model call: they need real workflow integration, strong engineering, human-centered automation, and feedback loops.

  • Systems, not isolated screens
  • Workflow-connected AI
  • Discovery to production
  • Continuous feedback
Visit Heru Loop
06 · Continuous evolution

Learning feeds the systems I build.

Formal study, production incidents, technical reading, teaching, and deliberate practice become better engineering judgment.

Formal education

MSc in Artificial Intelligence

Woolf University · LLMs, machine-learning systems, scalable AI, and applied AI engineering.

Current themes
  • Agentic AI systems
  • Distributed systems
  • Data platforms
  • Reliability engineering
  • Architecture and technical leadership

SystemCraft turns these themes into public technical education.

Operating agreementTerms & conditions apply in production

Principles I use when the system gets difficult.

  1. §1Systems, not screens
  2. §2Understand the problem before choosing the technology
  3. §3Reliability is a product feature
  4. §4Architecture must survive contact with production
  5. §5Make complexity visible
  6. §6Automate repeated engineering work
  7. §7Use AI to amplify judgment, not replace it
  8. §8Own the outcome, not only the ticket

Agreed — Ibrahim

Technical capabilities · grouped by problem

Open the relevant file.

01Distributed systems+

Systems that remain dependable as traffic, locations, teams, and failure modes grow.

Tools and methods

Event-driven architecture · Messaging · Coordination · Caching · Retries · Delivery guarantees

Where I used it

Applied across programmatic media delivery, synchronized playback, offline continuity, and fault-tolerant cloud services.

02Backend and APIs+

Service boundaries, APIs, and processing workflows designed for clear ownership and predictable production behavior.

Tools and methods

TypeScript, Node.js, NestJS, Go, Python, Java, .NET · REST, GraphQL, gRPC · RabbitMQ, Kafka

Where I used it

Built high-throughput processing, payment and subscription services, enterprise integrations, and reusable backend foundations.

03AI and agentic systems+

Reliable AI-enabled workflows connected to real knowledge, users, and operational constraints—not isolated model demos.

Tools and methods

LLM integration · RAG · Embeddings and retrieval · Agent memory and tools · Evaluation · Human-in-the-loop design

Where I used it

Applied to text-to-speech advertising, campaign recommendations, enterprise knowledge retrieval, and engineering documentation.

04Data platforms+

Data movement and analytical models designed around query patterns, freshness, cost, and production operability.

Tools and methods

PostgreSQL, MongoDB, Redis, Elasticsearch · BigQuery, Airflow, Spark, PySpark · Batch and streaming pipelines

Where I used it

Led a zero-downtime analytics modernization that reduced query latency by 70%, alongside enterprise ETL and reporting platforms.

05Frontend and product engineering+

Product workflows built end to end, with enough frontend depth to connect system behavior to the user experience.

Tools and methods

Angular, React, Next.js, React Native · WebSockets · Progressive web applications · Accessibility and responsive design

Where I used it

Delivered web and mobile products, data-intensive dashboards, reusable design-system components, and real-time customer workflows.

06Cloud and DevOps+

Delivery systems and observability treated as part of the product—not cleanup after implementation.

Tools and methods

AWS, GCP, Azure · Docker, Kubernetes, Terraform · CI/CD · OpenTelemetry, Prometheus, Loki · Incident response

Where I used it

Owned serverless media pipelines, multi-cloud analytics, standardized delivery across 40+ repositories, and long-term reliability improvements.

07Reliability and observability+

Production behavior made visible enough to diagnose failures, learn from incidents, and improve the system deliberately.

Tools and methods

OpenTelemetry · Prometheus · Loki · Structured logging · Alerting · Incident analysis · Performance profiling

Where I used it

Improved playback continuity, removed memory leaks, diagnosed distributed-service bottlenecks, and built observable delivery workflows.

08Architecture and technical strategy+

Business goals and constraints translated into system boundaries, explicit decisions, migration paths, and executable plans.

Tools and methods

Technical discovery · RFCs · ADRs · Architecture reviews · Migration planning · Roadmaps · Risk management

Where I used it

Owned long-term platform direction, customer solution design, zero-downtime migrations, and cross-team architecture decisions.

09Engineering leadership+

Teams supported with clear context, useful feedback, strong technical standards, and ownership of outcomes.

Tools and methods

Cross-functional squads · Mentoring · 1:1s · Performance reviews · Promotion cycles · Delivery leadership

Where I used it

Led and mentored international teams of 10+ engineers while remaining hands-on in architecture, code, and production.

Portrait of Ibrahim Elmansoury
Available worldwide
Late-night systems counsel · No billable hour required

Got a difficult system? Ibrahim can help make the case.

Ibrahim is open to Principal Engineer, Solution Architect, Staff Platform, and technical-leadership opportunities worldwide—including roles offering visa sponsorship or relocation.

Book 30 minutes ↗
ibrahim.elmansoury@gmail.com
LinkedIn ↗GitHub ↗View résumé ↗