Your event-driven AWS stack, reviewed by someone who runs one.
A fixed-price, senior audit of your EventBridge, SQS, Step Functions, Lambda and ECS workload, including the LLM calls running through it. Ten business days. You get what will break, what is costing you money, and the order to fix it in.
Is this for you
You are a funded startup spending USD 10-100K a month on AWS. Your backend is serverless and event-driven. You have 3-20 engineers and nobody whose whole job is the platform.
Something like this happened recently:
- The bill grew faster than traffic and nobody can say why.
- A retry storm or a full DLQ turned into an incident.
- Step Functions cost surprised someone.
- Due diligence or SOC 2 prep asked “who reviewed this?”
- The person who designed the stack left.
- Coding agents now write more of your Lambdas and Terraform than anyone has time to review.
- You put model calls into a queue or a workflow, and the bill or the error rate moved.
Not for you if you run on Kubernetes or EC2 fleets, if you have no production traffic yet, or if you want someone to do the fixes rather than find them. Also not for anything in B2B travel or airline software; I have a conflict there and will say so on the call.
What you get
A written report of 20-40 pages. One-page executive summary with the five findings that matter. Architecture as it actually exists in the account, reconstructed from the resources and IAM, not from your docs. Findings across eleven core areas, plus AI workloads when the stack calls models, each with evidence, severity, impact, and the fix. Cost analysis with numbers you can put in a spreadsheet.
A prioritized remediation backlog. Every finding as a ticket-ready item with effort, impact, dependencies, and suggested order. Markdown and CSV, so it imports into Linear or Jira.
A recorded walkthrough of 30-45 minutes, so the whole team can watch without another meeting.
A 60-minute Q&A call within five business days of delivery.
30 days of email follow-up for questions on any finding.
Read the sample report, an audit of a published AWS reference application.
What gets reviewed
| Area | What I look at |
|---|---|
| Event flow | EventBridge rules, SNS/SQS fan-out, ordering, duplicates, idempotency, schema drift |
| Failure handling | Retries, backoff, DLQs and whether anyone reads them, poison messages, partial batch failures |
| Step Functions | Standard vs Express fit, transition cost, polling patterns, Lambda glue that should be direct integrations |
| Lambda | Memory and timeout tuning, concurrency, provisioned concurrency spend, cold starts, VPC, runtime EOL |
| ECS / Fargate | Task sizing, autoscaling signals, deploy safety, Spot, log volume |
| Data | DynamoDB capacity mode, hot partitions, GSI cost, TTL, streams; S3 lifecycle |
| API edge | REST vs HTTP API, authorizers, throttling, caching, validation |
| Observability | Structured logs, retention cost, tracing coverage, alarms that would actually page |
| Security & IAM | Over-broad roles, resource policies, secrets, public endpoints |
| Cost | Top drivers, idle resources, NAT and data transfer, commitments fit |
| Delivery | IaC coverage and drift, environment parity, rollback path |
| AI workloads | LLM and Bedrock calls in queues and workflows: idempotency and retry cost, provider throttling and fallbacks, timeouts, token-spend alarms, what ends up in logs |
Every report answers five questions: what single failure takes the main flow down and would anyone be paged; where a duplicate or out-of-order event corrupts state; the three largest avoidable line items with numbers; what a compromised function can do; and whether the team can deploy a fix at 2 a.m. and roll it back.
How it works
| When | What |
|---|---|
| Day 0 | 20-minute fit call. Deposit. Start date locked. |
| Day 1 | 45-minute kickoff. You deploy a read-only IAM role from a template I provide and grant read access to your infra repo. |
| Days 1-6 | I collect evidence and analyze. Async questions only; about two hours of one engineer's time in total. |
| Day 7 | Mid-point note with the top findings so far. Nothing in the final report is a surprise. |
| Day 10 | Report and backlog delivered. |
| By day 15 | Walkthrough video and Q&A call. |
| +30 days | Follow-up window closes. |
Ten business days start at kickoff, not at payment.
Access
Read-only. One IAM role with ReadOnlyAccess plus CloudWatch Logs, X-Ray, and Cost Explorer read, locked to an external ID and a session limit. I never hold write permissions. You revoke the role the day the report lands. All evidence is deleted 30 days after the follow-up window closes.
Price
USD 3,500, fixed. No hourly billing. No change orders for one production workload in up to two accounts, roughly 50 Lambda functions, 10 state machines, 20 queues, topics or rules, and 10 DynamoDB tables. Bigger stacks get a custom quote on the fit call.
50% to book, 50% on delivery of the report.
Case-study rate: USD 2,100 for the first three audits booked before 30 November 2026, in exchange for a named case study and a testimonial.
Invoiced in USD from Argentina. Wire or Wise. W-8BEN available for US finance teams.
Who is doing the work
Steve Mallen. Fifteen years building backend systems, including five at Amazon, where I built a real-time security event platform for Ring from zero on ECS, SQS, Lambda and DynamoDB. Today I run event-driven serverless in production for a US B2B company at exactly this scale, including LLM-backed checks on the signup path, and I still carry the pager. Before that, Staff Engineer and tech lead.
You get the person who has been woken up by these failure modes, not a partner firm's junior with a checklist.
FAQ
Is this a Well-Architected Review? No. Partner-run WARs are broader, shallower, and can unlock AWS remediation credits; this cannot. If you need the credits, get the free partner WAR. If you need someone to find the retry storm before it finds you, this is that.
Will you do the fixes? Separately. The top three findings as a fixed two-week package for USD 8,000, or USD 1,500 a day with a five-day minimum for anything open-ended. Decide after you read the report.
Do you need production access? Read-only, yes. Metrics, logs, cost data and resource configuration are the evidence. Without them it is an opinion, not an audit.
What about our code? I read handlers to understand the event flow and check patterns like idempotency and retries. I do not review your domain logic. If coding agents wrote a lot of it, say so: I look harder at what they tend to get wrong, like swallowed errors, retries without idempotency and over-broad IAM.
Do you use AI to do the audit? Yes, for evidence collection and to sweep every resource instead of a sample. That is how one engineer delivers this in ten days. Every finding in the report is checked by me against evidence from your account. The audit role cannot read your secrets. Evidence only goes through model APIs that do not train on it, and log excerpts are masked first. If your data policy rules out third-party processing, say so on the call and it goes in the SOW before kickoff.
We don't use LLMs. Then that area is skipped and the other eleven get the time.
We use Terraform / SST / CDK / Serverless Framework. All fine.
Multi-region? Multiple workloads? Custom quote. Ask on the call.