Turn an unclear business problem into explicit goals, constraints, risks, and measurable outcomes.
Principal Software Engineer Solution Architect · Technical Leader
Ibrahim Elmansoury
Systems that endure. Ideas that move.
Ibrahim architects distributed platforms, intelligent products, and engineering teams that turn complicated business problems into dependable systems.
Based in Egypt · Working worldwide · Open to relocation
How complicated systems become dependable.
By Ibrahim Elmansoury · Principal Engineer
Impact across systems, AI products, and teams.
- 10+ yrs
- building production softwareArchitecture, implementation, delivery, and long-term system evolution
- AI → prod
- intelligent products deliveredFrom workflow discovery and model integration to dependable production systems
- RAG
- retrieval systems refinedImproving grounding, retrieval quality, evaluation, and operational feedback loops
- 10+
- engineers led and mentoredInternational cross-functional squads, 1:1s, performance reviews, and promotion cycles
- 100M+
- daily ad impressionsProgrammatic media delivery at retail scale
- 99.9%
- programmatic delivery successReliable campaign delivery across a distributed retail network
- 20K+
- connected media playersOperating across 3,000+ retail locations
- 3K+
- retail locationsA distributed physical media network
- 98%+
- playback reliabilityReliable delivery under variable connectivity
- <100 ms
- synchronization driftCoordinated multi-player audio and video
- ~90%
- faster report generationBackground processing and notification workflows
- 70%
- lower query latencyAfter a zero-downtime analytics migration
- 40+
- repositories standardizedUnified tooling, libraries, CI/CD, and delivery practices
- 1,387
- problems solvedDeliberate LeetCode practice begun during the COVID lockdown period
Classification · Principal-level engineering
The case for an engineer who works across the whole system.
Ibrahim moves between architecture, implementation, delivery, and leadership without treating them as separate jobs.
Outcomes over theatre.
- 01Architects distributed and cloud-native systems
- 02Works deeply in critical-path code
- 03Leads cross-functional international teams
- 04Turns ambiguity into reliable products
- 05Improves developer experience and delivery systems
- 06Builds across AI, data, backend, frontend, and platforms
- 07Mentors engineers and raises technical standards
What was happening, why it was difficult, how the solution worked, what changed, and the part I personally owned.
Enterprise advertising platform
Deliver programmatic campaigns reliably across a large, distributed physical network without sacrificing product velocity or operational control.
Campaign decisions had to reach thousands of physical devices, tolerate unreliable connectivity, and remain observable when delivery failed.
I led the platform architecture and long-term technical roadmap, designing cloud-native services around scalable delivery, fault tolerance, observability, and clear service boundaries.
The platform serves 100M+ daily impressions across 20,000+ players in 3,000+ retail locations with 99.9% programmatic delivery success and 98%+ playback reliability.
- Delivery guarantees over optimistic sends
- Operational visibility across a physical network
- Scale without coupling campaign logic to devices
More delivery state and operational discipline in exchange for reliability at retail scale.
I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.
In-store video synchronization
Play the same video across multiple screens inside a store while networks are unreliable and large media assets are expensive to move repeatedly.
Every screen had its own clock and buffering behavior. The design had to correct drift continuously, survive a lost connection, and avoid downloading the same large video repeatedly.
I designed scheduled playback epochs, WebSocket timing signals, local media prefetching, player-clock telemetry, continuous drift correction, and offline playback from the store edge.
Video playback across in-store screens stays below 100 ms drift, while adaptive delivery and edge caching reduced bandwidth consumption by approximately 60%.
- WebSocket timing signals coordinate a shared playback epoch
- Edge prefetching trades local storage for lower bandwidth
- Clock telemetry drives correction without visibly restarting video
Local device complexity in exchange for continuity, synchronization, and lower bandwidth.
I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.
Analytics workflow modernization
Synchronous report generation blocked user workflows and became too slow as analytical volume and reporting complexity grew.
The workflow had to execute expensive analytical queries and generate large files without holding an HTTP request open, duplicating jobs, losing progress, or hiding failures.
I designed a durable asynchronous workflow that validates requests, persists job state, queues work, processes partitions in parallel, queries analytical stores, streams generated artifacts into object storage, and notifies users when results are ready. Retries, dead-letter handling, and job-level observability make failures recoverable.
Report-generation time fell by approximately 90%, with a clearer and more resilient user workflow.
- 02Persist jobstatus · owner · filters
- 03Durable queuepriority · retry policy
- 04Worker poolparallel partitions
- 05Query layerwarehouse · cache
- 06Build artifactstream · compress
- 07Object storagesigned result URL
- 08Notify userin-app · email
- Persist job state before queueing expensive work
- Parallel workers generate and stream artifacts without holding the request
- Retries, dead letters, and notifications make failure recoverable
More workflow and job-state machinery in exchange for a faster, more resilient reporting experience.
I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.
Data platform re-architecture
Operational and analytical workloads competed for resources, producing slow queries while the business required uninterrupted service.
The migration had to separate analytical work from production traffic without downtime or a risky one-time cutover.
I re-architected the analytical path and led a zero-downtime migration using Airflow, PySpark, BigQuery, TimescaleDB, and materialized views.
Analytical query latency fell by approximately 70% without interrupting production.
- Separate operational and analytical workloads
- Migrate without a business-visible cutover
- Design orchestration for repeatability and recovery
Migration and pipeline complexity in exchange for faster analytics and safer evolution.
I owned the architecture and technical direction, worked hands-on in the critical paths, coordinated delivery, and used production feedback to improve the system.
Two product concepts that combine generative AI with professional controls, transparent reasoning, and dependable system design.
Photo direction through conversation.
A professional image workspace where a user can describe intent in natural language, compare revisions, and make precise changes directly on the canvas.
Creative teams could generate images quickly, but prompt-only tools made precise direction, repeatable revisions, and brand-safe review difficult.
Natural-language intent had to coexist with canvas-level control, protected product details, asynchronous generation, and a reviewable history.
A conversational workspace that turns requests into structured edits while masks, versioning, approval, and export remain explicit.
Teams move faster without surrendering the precision, traceability, and human approval required for professional creative work.
Premium studio
The conversation becomes structured edit instructions. The canvas applies them non-destructively, preserves locked regions, and keeps every revision reproducible.
Ibrahim shaped the product model, interaction between conversation and canvas, AI orchestration boundaries, production controls, and human approval path.
Model orchestration, asset storage, safety checks, version history, asynchronous generation, and human approval add machinery—but turn a prompt box into a dependable workflow.
Find the company before the vacancy.
An intelligence workflow that discovers relevant companies, understands their direction, monitors suitable roles, and explains why each opportunity deserves attention.
Traditional job search started with published vacancies and title keywords, missing relevant companies before the right role appeared.
Company direction, role meaning, engineering culture, mobility constraints, source freshness, and personal goals all had to be reasoned about together.
An explainable discovery pipeline that builds a career profile, finds companies, monitors role signals, ranks fit, and exposes the evidence behind every match.
The user receives a focused, continuously refreshed shortlist—and stays responsible for every save, dismissal, and final decision.
Northstar Labs
Strong architecture ownership and an international platform mandate.
- ✓ Distributed systems
- ✓ Technical leadership
- ✓ Relocation
Role description · company trajectory · engineering culture · location policy · Ibrahim’s experience
The engine starts with professional intent, then discovers companies and roles through meaning, trajectory, and constraints—not title matching alone.
Ibrahim designed the discovery workflow, ranking architecture, explainability model, system boundaries, and feedback controls that keep people responsible for the shortlist.
Source ingestion, entity resolution, embeddings, search, LLM reasoning, freshness checks, and deduplication add complexity in exchange for explainable, current recommendations.
I move from business ambiguity into implementation, production, and team leverage—without floating above the work.
Implementation
- APIs and distributed services
- Data models and pipelines
- Critical code and debugging
Consultant, technical lead, senior engineer, principal engineer, solution architect, and Lead Solutions Architect—each role expands the scope while preserving hands-on ownership.
01Heru LoopLead Solutions Architect
I lead the technical direction of an IT services and product-engineering company that helps organisations move from an ambiguous idea to reliable production software.
Shape product and technology strategy across digital platforms, AI-enabled systems, workflow automation, and scalable custom software.
Lead technical discovery with clients, translating business constraints into architecture, phased delivery plans, and explicit engineering trade-offs.
Remain hands-on in system design, implementation decisions, code review, infrastructure, and the critical paths that carry the most delivery risk.
Read details +
Lead Solutions Architect
I lead the technical direction of an IT services and product-engineering company that helps organisations move from an ambiguous idea to reliable production software.
Shape product and technology strategy across digital platforms, AI-enabled systems, workflow automation, and scalable custom software.
Lead technical discovery with clients, translating business constraints into architecture, phased delivery plans, and explicit engineering trade-offs.
Remain hands-on in system design, implementation decisions, code review, infrastructure, and the critical paths that carry the most delivery risk.
02Intouch.comPrincipal Software Engineer · Solution Architect
I own technical direction for high-scale media, analytics, and AI-enabled engineering systems while leading a cross-functional team of 10+ engineers.
Own the technical vision and long-term evolution of a programmatic media platform serving 100M+ daily impressions.
Lead and mentor 10+ engineers through architecture reviews, design guidance, incident analysis, and continuous technical feedback.
Mentor engineers across cross-functional squads through regular 1:1s, performance reviews, promotion cycles, growth plans, and clear technical feedback.
Read details +
Principal Software Engineer · Solution Architect
I own technical direction for high-scale media, analytics, and AI-enabled engineering systems while leading a cross-functional team of 10+ engineers.
Own the technical vision and long-term evolution of a programmatic media platform serving 100M+ daily impressions.
Lead and mentor 10+ engineers through architecture reviews, design guidance, incident analysis, and continuous technical feedback.
Mentor engineers across cross-functional squads through regular 1:1s, performance reviews, promotion cycles, growth plans, and clear technical feedback.
03Intouch.comSenior Software Engineer
I built high-throughput backend and data systems, improved production reliability, and helped modernize the platform around reusable services and engineering standards.
Built a .NET batch-processing system that reduced a 1M+ record import from three hours to 25 minutes.
Developed backend platforms for campaign management, media delivery, analytics, reporting, and enterprise workflows.
Built Airflow-orchestrated ETL pipelines across cloud analytics platforms.
Read details +
Senior Software Engineer
I built high-throughput backend and data systems, improved production reliability, and helped modernize the platform around reusable services and engineering standards.
Built a .NET batch-processing system that reduced a 1M+ record import from three hours to 25 minutes.
Developed backend platforms for campaign management, media delivery, analytics, reporting, and enterprise workflows.
Built Airflow-orchestrated ETL pipelines across cloud analytics platforms.
04HeartAttackTech Lead · Individual Contributor
I led web, mobile, and backend delivery while continuing to design and implement the systems behind customer products.
Owned technical delivery from solution design and planning through production deployment and operational support.
Coordinated frontend, backend, and mobile engineers while managing dependencies, release risks, and incremental delivery.
Established reusable components, shared libraries, testing practices, documentation, Git workflows, and CI/CD standards.
Read details +
Tech Lead · Individual Contributor
I led web, mobile, and backend delivery while continuing to design and implement the systems behind customer products.
Owned technical delivery from solution design and planning through production deployment and operational support.
Coordinated frontend, backend, and mobile engineers while managing dependencies, release risks, and incremental delivery.
Established reusable components, shared libraries, testing practices, documentation, Git workflows, and CI/CD standards.
05IndependentSoftware Engineering Consultant
I partnered with startups and growing businesses from discovery through production across commerce, logistics, mobility, and enterprise operations.
Led customer discovery, architecture workshops, technical planning, implementation, deployment, and production support.
Modernized legacy platforms through cloud adoption, API integrations, clearer service boundaries, and performance optimization.
Improved production systems using database tuning, asynchronous processing, caching, validation, monitoring, and error handling.
Read details +
Software Engineering Consultant
I partnered with startups and growing businesses from discovery through production across commerce, logistics, mobility, and enterprise operations.
Led customer discovery, architecture workshops, technical planning, implementation, deployment, and production support.
Modernized legacy platforms through cloud adoption, API integrations, clearer service boundaries, and performance optimization.
Improved production systems using database tuning, asynchronous processing, caching, validation, monitoring, and error handling.
Education, developer tooling, and AI-era software engineering—each has a distinct problem, audience, and operating model.

SystemCraft
An interactive learning platform that makes distributed systems and software architecture approachable through scenarios, trade-offs, visual explanations, and hands-on exploration.
Software engineers often learn patterns by name before they learn when those patterns help, fail, or create new trade-offs.
- Distributed systems
- Consistency
- Time and ordering
- Sagas
- Reliability
- Architecture trade-offs
ShikamaruOne command · many moving partsShikamaru
Developer tooling that turns repetitive engineering workflows into dependable automation across repositories, local infrastructure, and Azure DevOps.
Engineers lose attention switching between repositories, environment setup, service dependencies, and repetitive platform operations.
- Multi-repo automation
- Docker environments
- Azure DevOps synchronization
Human-ledJudgmentsets intentand approves outcomes
AI accelerates each stage. People remain accountable for the direction.
Heru Loop
An AI-native software studio building intelligent workflows, agentic systems, living products, and the reliable platforms behind them.
Useful AI products need more than a model call: they need real workflow integration, strong engineering, human-centered automation, and feedback loops.
- Systems, not isolated screens
- Workflow-connected AI
- Discovery to production
- Continuous feedback
Engineering leverage that compounds.
I publish tools and learning material when repeated friction reveals a problem worth solving for other engineers.
Shikamaru CLI
A command-line toolkit for automating multi-repository environments, Docker-based infrastructure setup, and Azure DevOps synchronization.
I built it to reduce the operational and cognitive overhead that accumulates when developers work across many services and environments.Reusable engineering systems
Shared libraries, backend frameworks, delivery pipelines, observability patterns, and development standards.
Built to reduce duplicated decisions and focus teams on product-specific problems.How I spent the quiet part of COVID.
I worked systematically through LeetCode to strengthen how I decompose unfamiliar problems, test assumptions, reason about trade-offs, and translate an approach into efficient code.
1,387 problems solved. The habit outlasted lockdown: stay patient with ambiguity and keep improving the model.
Open my LeetCode profile ↗- 291
- Easy
- 797
- Medium
- 299
- Hard
Most recent · 365 Days Badge
Learning feeds the systems I build.
Formal study, production incidents, technical reading, teaching, and deliberate practice become better engineering judgment.
MSc in Artificial Intelligence
Woolf University · LLMs, machine-learning systems, scalable AI, and applied AI engineering.
- Agentic AI systems
- Distributed systems
- Data platforms
- Reliability engineering
- Architecture and technical leadership
SystemCraft turns these themes into public technical education.
Principles I use when the system gets difficult.
- §1Systems, not screens
- §2Understand the problem before choosing the technology
- §3Reliability is a product feature
- §4Architecture must survive contact with production
- §5Make complexity visible
- §6Automate repeated engineering work
- §7Use AI to amplify judgment, not replace it
- §8Own the outcome, not only the ticket
Agreed — Ibrahim
Open the relevant file.
01Distributed systems+
Systems that remain dependable as traffic, locations, teams, and failure modes grow.
Tools and methodsEvent-driven architecture · Messaging · Coordination · Caching · Retries · Delivery guarantees
Where I used itApplied across programmatic media delivery, synchronized playback, offline continuity, and fault-tolerant cloud services.
02Backend and APIs+
Service boundaries, APIs, and processing workflows designed for clear ownership and predictable production behavior.
Tools and methodsTypeScript, Node.js, NestJS, Go, Python, Java, .NET · REST, GraphQL, gRPC · RabbitMQ, Kafka
Where I used itBuilt high-throughput processing, payment and subscription services, enterprise integrations, and reusable backend foundations.
03AI and agentic systems+
Reliable AI-enabled workflows connected to real knowledge, users, and operational constraints—not isolated model demos.
Tools and methodsLLM integration · RAG · Embeddings and retrieval · Agent memory and tools · Evaluation · Human-in-the-loop design
Where I used itApplied to text-to-speech advertising, campaign recommendations, enterprise knowledge retrieval, and engineering documentation.
04Data platforms+
Data movement and analytical models designed around query patterns, freshness, cost, and production operability.
Tools and methodsPostgreSQL, MongoDB, Redis, Elasticsearch · BigQuery, Airflow, Spark, PySpark · Batch and streaming pipelines
Where I used itLed a zero-downtime analytics modernization that reduced query latency by 70%, alongside enterprise ETL and reporting platforms.
05Frontend and product engineering+
Product workflows built end to end, with enough frontend depth to connect system behavior to the user experience.
Tools and methodsAngular, React, Next.js, React Native · WebSockets · Progressive web applications · Accessibility and responsive design
Where I used itDelivered web and mobile products, data-intensive dashboards, reusable design-system components, and real-time customer workflows.
06Cloud and DevOps+
Delivery systems and observability treated as part of the product—not cleanup after implementation.
Tools and methodsAWS, GCP, Azure · Docker, Kubernetes, Terraform · CI/CD · OpenTelemetry, Prometheus, Loki · Incident response
Where I used itOwned serverless media pipelines, multi-cloud analytics, standardized delivery across 40+ repositories, and long-term reliability improvements.
07Reliability and observability+
Production behavior made visible enough to diagnose failures, learn from incidents, and improve the system deliberately.
Tools and methodsOpenTelemetry · Prometheus · Loki · Structured logging · Alerting · Incident analysis · Performance profiling
Where I used itImproved playback continuity, removed memory leaks, diagnosed distributed-service bottlenecks, and built observable delivery workflows.
08Architecture and technical strategy+
Business goals and constraints translated into system boundaries, explicit decisions, migration paths, and executable plans.
Tools and methodsTechnical discovery · RFCs · ADRs · Architecture reviews · Migration planning · Roadmaps · Risk management
Where I used itOwned long-term platform direction, customer solution design, zero-downtime migrations, and cross-team architecture decisions.
09Engineering leadership+
Teams supported with clear context, useful feedback, strong technical standards, and ownership of outcomes.
Tools and methodsCross-functional squads · Mentoring · 1:1s · Performance reviews · Promotion cycles · Delivery leadership
Where I used itLed and mentored international teams of 10+ engineers while remaining hands-on in architecture, code, and production.

Got a difficult system? Ibrahim can help make the case.
Ibrahim is open to Principal Engineer, Solution Architect, Staff Platform, and technical-leadership opportunities worldwide—including roles offering visa sponsorship or relocation.