Your event-driven AWS stack, reviewed by someone who runs one.

A fixed-price, senior audit of your EventBridge, SQS, Step Functions, Lambda and ECS workload, including the LLM calls running through it. Ten business days. You get what will break, what is costing you money, and the order to fix it in.

USD 3,500 fixed10 business daysRead-only accessNo upsell pressure

Book a fit call

Is this for you

You are a funded startup spending USD 10-100K a month on AWS. Your backend is serverless and event-driven. You have 3-20 engineers and nobody whose whole job is the platform.

Something like this happened recently:

Not for you if you run on Kubernetes or EC2 fleets, if you have no production traffic yet, or if you want someone to do the fixes rather than find them. Also not for anything in B2B travel or airline software; I have a conflict there and will say so on the call.

What you get

A written report of 20-40 pages. One-page executive summary with the five findings that matter. Architecture as it actually exists in the account, reconstructed from the resources and IAM, not from your docs. Findings across eleven core areas, plus AI workloads when the stack calls models, each with evidence, severity, impact, and the fix. Cost analysis with numbers you can put in a spreadsheet.

A prioritized remediation backlog. Every finding as a ticket-ready item with effort, impact, dependencies, and suggested order. Markdown and CSV, so it imports into Linear or Jira.

A recorded walkthrough of 30-45 minutes, so the whole team can watch without another meeting.

A 60-minute Q&A call within five business days of delivery.

30 days of email follow-up for questions on any finding.

Read the sample report, an audit of a published AWS reference application.

What gets reviewed

AreaWhat I look at
Event flowEventBridge rules, SNS/SQS fan-out, ordering, duplicates, idempotency, schema drift
Failure handlingRetries, backoff, DLQs and whether anyone reads them, poison messages, partial batch failures
Step FunctionsStandard vs Express fit, transition cost, polling patterns, Lambda glue that should be direct integrations
LambdaMemory and timeout tuning, concurrency, provisioned concurrency spend, cold starts, VPC, runtime EOL
ECS / FargateTask sizing, autoscaling signals, deploy safety, Spot, log volume
DataDynamoDB capacity mode, hot partitions, GSI cost, TTL, streams; S3 lifecycle
API edgeREST vs HTTP API, authorizers, throttling, caching, validation
ObservabilityStructured logs, retention cost, tracing coverage, alarms that would actually page
Security & IAMOver-broad roles, resource policies, secrets, public endpoints
CostTop drivers, idle resources, NAT and data transfer, commitments fit
DeliveryIaC coverage and drift, environment parity, rollback path
AI workloadsLLM and Bedrock calls in queues and workflows: idempotency and retry cost, provider throttling and fallbacks, timeouts, token-spend alarms, what ends up in logs

Every report answers five questions: what single failure takes the main flow down and would anyone be paged; where a duplicate or out-of-order event corrupts state; the three largest avoidable line items with numbers; what a compromised function can do; and whether the team can deploy a fix at 2 a.m. and roll it back.

How it works

WhenWhat
Day 020-minute fit call. Deposit. Start date locked.
Day 145-minute kickoff. You deploy a read-only IAM role from a template I provide and grant read access to your infra repo.
Days 1-6I collect evidence and analyze. Async questions only; about two hours of one engineer's time in total.
Day 7Mid-point note with the top findings so far. Nothing in the final report is a surprise.
Day 10Report and backlog delivered.
By day 15Walkthrough video and Q&A call.
+30 daysFollow-up window closes.

Ten business days start at kickoff, not at payment.

Access

Read-only. One IAM role with ReadOnlyAccess plus CloudWatch Logs, X-Ray, and Cost Explorer read, locked to an external ID and a session limit. I never hold write permissions. You revoke the role the day the report lands. All evidence is deleted 30 days after the follow-up window closes.

Price

USD 3,500, fixed. No hourly billing. No change orders for one production workload in up to two accounts, roughly 50 Lambda functions, 10 state machines, 20 queues, topics or rules, and 10 DynamoDB tables. Bigger stacks get a custom quote on the fit call.

50% to book, 50% on delivery of the report.

Guarantee. If you do not think the report was worth the fee, the second half is waived. No conditions.

Case-study rate: USD 2,100 for the first three audits booked before 30 November 2026, in exchange for a named case study and a testimonial.

Invoiced in USD from Argentina. Wire or Wise. W-8BEN available for US finance teams.

Who is doing the work

Steve Mallen. Fifteen years building backend systems, including five at Amazon, where I built a real-time security event platform for Ring from zero on ECS, SQS, Lambda and DynamoDB. Today I run event-driven serverless in production for a US B2B company at exactly this scale, including LLM-backed checks on the signup path, and I still carry the pager. Before that, Staff Engineer and tech lead.

You get the person who has been woken up by these failure modes, not a partner firm's junior with a checklist.

FAQ

Is this a Well-Architected Review? No. Partner-run WARs are broader, shallower, and can unlock AWS remediation credits; this cannot. If you need the credits, get the free partner WAR. If you need someone to find the retry storm before it finds you, this is that.

Will you do the fixes? Separately. The top three findings as a fixed two-week package for USD 8,000, or USD 1,500 a day with a five-day minimum for anything open-ended. Decide after you read the report.

Do you need production access? Read-only, yes. Metrics, logs, cost data and resource configuration are the evidence. Without them it is an opinion, not an audit.

What about our code? I read handlers to understand the event flow and check patterns like idempotency and retries. I do not review your domain logic. If coding agents wrote a lot of it, say so: I look harder at what they tend to get wrong, like swallowed errors, retries without idempotency and over-broad IAM.

Do you use AI to do the audit? Yes, for evidence collection and to sweep every resource instead of a sample. That is how one engineer delivers this in ten days. Every finding in the report is checked by me against evidence from your account. The audit role cannot read your secrets. Evidence only goes through model APIs that do not train on it, and log excerpts are masked first. If your data policy rules out third-party processing, say so on the call and it goes in the SOW before kickoff.

We don't use LLMs. Then that area is skipped and the other eleven get the time.

We use Terraform / SST / CDK / Serverless Framework. All fine.

Multi-region? Multiple workloads? Custom quote. Ask on the call.

Book a fit call