Sample audit report

The audit as a client receives it, run against a published AWS reference application instead of a client's stack.

Version 0.9: static review complete. All 24 findings, the cost model and the backlog are final. The 41 items marked pending live run need metrics from a deployed copy and are filled in when that run lands. The offer · Book a fit call

Client: AnyCompany (sample engagement; subject is the open-source aws-samples/app-2025 application)
Workload: Order and transaction event pipeline
Accounts: one sandbox account, one region
Engagement dates: 2026-09-12 (static review) — Live run: deploy date
Prepared by: Steve Mallen — stvmallen.com
Version: 0.9 — static review complete; live evidence pending

This is the public sample report for the audit offer at stvmallen.com/audit. The subject is a real, published AWS reference application, reviewed exactly as a client stack would be. Where a finding needs metrics from a running deployment, it says so.


1. Executive summary

In one paragraph. The pipeline is well shaped: a single custom EventBridge bus, Step Functions consuming events through direct service integrations, SNS/SQS for the legacy fan-out, and Firehose for the analytical copy. It is also unattended: there is no dead-letter queue, retry policy, alarm, or trace anywhere in the stack, one of the two SQS queues accepts messages from any AWS account, the function that publishes every downstream event reports failure as success, and the highest-volume flow runs on a Standard workflow that would cost roughly sixty times what an Express workflow would at production rates.

The five findings that matter

# Finding Sev Impact Fix effort
F-03 The publish-events function returns an error object instead of throwing, so every workflow that publishes an event through it treats a failed PutEvents as success. Downstream events are silently lost. S1 Reliability: data loss S
F-04 LegacyOperationQueue policy uses ForAllValues:ArnEquals on the single-valued aws:SourceArn key with Principal: "*". A request with no source ARN passes the condition, so any AWS principal can send messages to the queue. S1 Security S
F-05 TransactionProcess, the per-transaction flow, is a Standard workflow with two states. At 10 transactions/s that is ~USD 1,950/month in state transitions; the same flow on Express is ~USD 30. S1 Cost ≈ USD 1,900/month at production rate S
F-06 No dead-letter queue on any EventBridge target, SQS queue, or Lambda event source mapping. A failed delivery or a poison message is either dropped or retried until the message expires. S1 Reliability M
F-01 All three functions run nodejs12.x (EOL 2022) with aws-sdk v2. AWS no longer allows creating or updating functions on this runtime; the stack cannot be redeployed. S1 Operability: cannot deploy a fix M

Numbers

Metric Value
Findings total 24 (S1: 6, S2: 9, S3: 6, Info: 3)
Identified monthly savings USD ~1,920 at 10 TPS steady; Live run: at measured rate
Current monthly AWS spend (workload) pending live run
Unmonitored failure modes found 9 (no alarms exist)
Resources outside IaC 1 known (simulate-transaction function, no template); Live run: anything else

What to do this month

  1. F-03 and F-04: two one-line fixes that close a data-loss bug and an open queue. Same afternoon.
  2. F-06 with F-09: add DLQs to every target and Retry/Catch to every task, then alarm on DLQ depth (F-17). This is the safety net the stack does not have.
  3. F-05: switch TransactionProcess to Express. Requires F-03 first, because Express executions do not retry at the workflow level and the publisher must fail loudly.

Answers to the five questions

  1. What single failure takes the main flow down, and would anyone be paged? A PutEvents throttle or outage. The publisher returns success anyway (F-03), the workflow ends green, and nothing pages because there are no alarms (F-17). Nobody would know.
  2. Where does a duplicate or out-of-order event corrupt state? Duplicates are safe on the transaction path: DynamoDB PutItem keyed on the EventBridge event id is idempotent (see §7). The account-normalization callback is not: a duplicate account-created event opens a second Standard execution that waits up to a year for a token nobody will send (F-10).
  3. The three largest avoidable line items. Standard workflow transitions (F-05, ~USD 1,900/month at 10 TPS); the raw S3 bucket with no lifecycle policy (F-14, grows without bound); Express workflow logging at ALL with execution data (F-19, Live run: ingestion volume).
  4. What can a compromised function do? Little. Roles are per-workflow and resource-scoped (§7). The exception is the queue policy (F-04), which is an external exposure rather than a lateral one.
  5. Can the team deploy a fix at 2 a.m. and roll it back? No. The runtime is deprecated (F-01), so sam deploy fails on the function resources until the migration in F-02 is done. Hard-coded resource names (F-22) also prevent a parallel stack for testing the fix.

2. Scope and method

In scope: the five SAM stacks in aws-samples/app-2025 at commit d3e3915 (2020-07-22): infrastructure, customer, operations, billing, simulator. One account, one region.
Out of scope: the Athena/QuickSight reporting layer described in the README (console-only, no IaC); the three Twitch-episode transcripts.
Access used: static review of templates and handler code. Live run: read-only role audit-ro with external ID, assumed dates.
Evidence window: Live run: 3 hours of simulator traffic at 5 TPS; no historical cost data exists for a fresh deployment.
Method: per-service checklist review against the IaC; handler-level code reading; cost modeling from published AWS pricing at an assumed steady 10 transactions/s, which is the simulator's default DESIRED_TPS.

Limitations. This is a small workload, at the bottom of the size band the audit is designed for, so the report is shorter than a typical one. To deploy it for live evidence the runtime and SDK had to be patched (see F-01, F-02); everything else was reviewed exactly as published. Cost figures are modeled, not billed, and are labeled with their assumptions.


3. Architecture as found

Architecture diagram: producers publish to the AnyCompany EventBridge bus; rules route to Firehose (S3, Glue), three Step Functions workflows (DynamoDB, publish-events Lambda, SQS callback), and an SNS topic feeding the legacy SQS queue.
Reconstructed from the templates. Click to open full size.

Main flows

Flow Entry Transport Consumers Volume/day p99 latency
Transaction transaction-initiated event EventBridge → Firehose + Standard SFN 2 pending live run (864K at 10 TPS) pending live run
Account normalization account-created event EventBridge → Standard SFN → SQS callback 1 pending live run pending live run
Subscription expiry subscription-expired event EventBridge → Express SFN 1 pending live run pending live run
Legacy operations operation-performed event EventBridge → SNS → SQS 1 (external) pending live run n/a

Where the diagram and the docs differ

Inventory summary

Resource Count In IaC Notes
Lambda functions 3 2 nodejs12.x; simulate-transaction has code but no template
State machines 3 3 2 Standard, 1 Express
Event buses / rules 1 / 4 1 / 4
SQS queues / SNS topics 2 / 1 2 / 1
DynamoDB tables 1 1 SimpleTable, on-demand
Firehose streams 1 1
S3 buckets 2 2
Glue databases / tables 1 / 1 1 / 1
Log groups 1 explicit + 3 implicit 1 Lambda groups are created at first invoke, no retention
ECS services 0
API Gateway APIs 0 no synchronous entry point; producers call PutEvents directly

4. Findings

4.1 Event flow

F-06 — No dead-letter queue on any EventBridge target

Severity: S1 Area: §1 Effort: M
Impact: Reliability. A target that fails after EventBridge's own retries (up to 24 h / 185 attempts) drops the event with no record.

What I found. None of the four rules' targets (billing/template.yaml:38-65, customer/template.yaml:451-454, 562-565, operations/template.yaml:848-850) sets DeadLetterConfig. The Firehose target and the three Step Functions targets fail if the destination throttles or the role is broken; the SNS target fails if the topic policy changes.

Evidence. Templates as cited. Live run: FailedInvocations per rule over the evidence window.

Why it matters. StartExecution on a Standard workflow is throttled at a bucket of 1,300 with a refill of 300/s in us-east-1 (lower elsewhere). A burst above that loses transactions after retries expire, and the Firehose copy and the DynamoDB copy will disagree with no way to reconcile.

Recommendation. One SQS DLQ per stack, DeadLetterConfig.Arn on every target, and a RetryPolicy with MaximumRetryAttempts and MaximumEventAgeInSeconds set explicitly rather than the defaults. Alarm on DLQ depth (F-17). Enable an archive on the bus for replay (F-08).

Savings / cost of fix. None; reliability only. DLQ cost is negligible.

F-07 — NormalizeAccountProcess sends the literal string "$.detail.crm_data" to the queue

Severity: S2 Area: §1 Effort: S
Impact: Reliability: the downstream consumer never receives the CRM data.

What I found. customer/template.yaml:483 reads "Message.?": "$.detail.crm_data". The ASL path-selection suffix is .$; .? is not an operator, so States Language treats it as a literal key and the message body carries the string $.detail.crm_data, not the object. CloudFormation accepts the definition because it is valid JSON.

Evidence. Template line cited. Live run: sample message body from AccountCreationQueue.

Why it matters. The simulated approver ignores the message body, so the bug is invisible in this repo. A real consumer would get a JSONPath string where it expects a customer record. This is the kind of typo a schema registry or a contract test catches and a demo does not.

Recommendation. "Message.$": "$.detail.crm_data". Add a contract test that starts the workflow with a fixture and asserts the queue message shape.

Savings / cost of fix. None.

F-08 — No event archive or schema registry on the bus

Severity: S3 Area: §1 Effort: S
Impact: Operability: no replay after a consumer bug, no contract for producers.

What I found. infrastructure/template.yaml:783-786 declares the bus with no AWS::Events::Archive and no schema discovery.

Recommendation. Add an archive with a retention matching the business need (30–90 days) and enable schema discovery on the bus for a week to capture the four event shapes, then commit them.

F-23 — LegacyOperationQueue has no consumer in IaC

Severity: S3 Area: §1, §11 Effort: S
Impact: Operability.

What I found. The queue is an output of the operations stack (operations/template.yaml:886-888) and nothing in the repository reads it. Whatever does is outside IaC. Live run: NumberOfMessagesReceived on the queue; if zero, the queue is dead and the rule can go.

Recommendation. Either bring the consumer into IaC or delete the rule, topic and queue.

I-02 — SNS topic with a single SQS subscriber

Severity: Info Area: §1

LegacyOperationTopic fans out to exactly one queue. EventBridge can target SQS directly; the SNS hop adds a delivery, a policy, and a failure point. Keep it only if a second subscriber is planned. Not a finding on its own.

4.2 Failure handling

F-03 — publish-events reports a failed PutEvents as success

Severity: S1 Area: §2 Effort: S
Impact: Reliability: silent loss of every downstream event (transaction-processed, account-normalized, expiration-processed).

What I found. infrastructure/publish-events/app.js:755-758: when result.FailedEntryCount > 0 the handler returns an object with an Error key. It does not throw. Step Functions' lambda:invoke integration sees a successful invocation with a payload, and all three workflows that call it (billing/template.yaml:224-245, customer/template.yaml:490-503, 613-626) mark the step complete. None of them inspects the payload.

Evidence. Code and template lines cited. Live run: force a PutEvents failure by revoking events:PutEvents on the role for one execution; the execution completes green.

Why it matters. PutEvents fails on throttling (the default quota is 10,000 requests/s in large regions but far lower in small ones), on payload size, and during EventBridge incidents. Every such failure is a transaction-processed event that never existed, with a green execution history saying it did.

Recommendation. Throw on FailedEntryCount > 0 (throw new Error(...)) so the invoke task fails and Retry (F-09) applies. Better: remove the function entirely and use the arn:aws:states:::events:putEvents direct integration, which fails the task natively and removes a runtime to maintain (F-11).

Savings / cost of fix. None; reliability only.

F-09 — No Retry or Catch on any task in any workflow

Severity: S1 Area: §2 Effort: M
Impact: Reliability. Any transient error fails the execution outright; there is no error path.

What I found. Zero Retry or Catch blocks across the three definitions (billing/template.yaml:200-247, customer/template.yaml:463-505, 580-628). The DynamoDB PutItem task, the SQS SendMessage task and the three Lambda invokes all fail the execution on the first throttle or 5xx.

Why it matters. DynamoDB on-demand tables throttle during ramp-up above the previous peak (F-15). Lambda invokes return Lambda.TooManyRequestsException under account concurrency pressure. Each becomes a FAILED execution nobody is alarmed on (F-17), with the transaction written to Firehose but not to DynamoDB.

Recommendation. A Retry block on every task with ErrorEquals: ["States.TaskFailed", "Lambda.ServiceException", "Lambda.TooManyRequestsException", "DynamoDB.ProvisionedThroughputExceededException"], IntervalSeconds: 2, BackoffRate: 2, MaxAttempts: 4, and a Catch on the last task that routes to a Fail state with a meaningful cause. Route failed executions to a DLQ via an EventBridge rule on Step Functions Execution Status Change → FAILED.

F-10 — Callback task has no timeout or heartbeat

Severity: S2 Area: §2, §3 Effort: S
Impact: Reliability and cost. Orphaned executions wait for up to one year.

What I found. customer/template.yaml:477-489: the sqs:sendMessage.waitForTaskToken task sets neither TimeoutSeconds nor HeartbeatSeconds. If the message is lost, the token expires, or the consumer fails (F-12), the execution stays RUNNING until the Standard workflow's one-year maximum.

Why it matters. Every orphan occupies one of the account's open-execution quota and, if the account-created event is ever duplicated, produces two executions, one of which never completes. Cleanup is manual.

Recommendation. TimeoutSeconds matched to the SLA for the human step (hours, not a year), a Catch on States.Timeout that routes to an escalation path, and idempotent execution names derived from the account id so a duplicate event is rejected by StartExecution instead of creating a twin.

F-12 — simulate-approval handles SQS batches incorrectly and has no partial-batch failure handling

Severity: S2 Area: §2, §4 Effort: S
Impact: Reliability: records silently dropped or infinitely retried.

What I found. simulator/simulate-approval/app.js:901-925 loops over event.Records but calls callback() inside the loop after the first asynchronous sendTaskSuccess resolves. With a batch size above one, the invocation ends after the first record; the rest are deleted from the queue unprocessed because Lambda treats the invocation as successful. Conversely, if any record's token is already expired, callback(err) fails the whole batch, and the event source mapping (simulator/template.yaml:1058-1062) has no FunctionResponseTypes: [ReportBatchItemFailures] and the queue has no redrive policy (F-13), so the entire batch is retried until the 4-day message retention expires.

Recommendation. Async handler with Promise.allSettled over records, return batchItemFailures for the ones that failed, enable ReportBatchItemFailures on the mapping, and add a redrive policy on the queue.

F-13 — SQS queues have no redrive policy and default visibility timeouts

Severity: S2 Area: §2 Effort: S
Impact: Reliability.

What I found. AccountCreationQueue (customer/template.yaml:544-546) and LegacyOperationQueue (operations/template.yaml:860-863) are declared with no properties: no RedrivePolicy, 30-second visibility timeout, 4-day retention, no long polling. AccountCreationQueue is consumed by a Lambda with a 3-second timeout, so the 6× rule holds by accident.

Recommendation. RedrivePolicy with maxReceiveCount: 5 to a DLQ per queue; ReceiveMessageWaitTimeSeconds: 20; MessageRetentionPeriod set deliberately; VisibilityTimeout explicitly at 6× consumer timeout.

4.3 Step Functions

F-05 — TransactionProcess runs as a Standard workflow

Severity: S1 Area: §3 Effort: S
Impact: Cost ≈ USD 1,900/month at 10 TPS; scales linearly with volume.

What I found. billing/template.yaml:193-252: StateMachineType: STANDARD for a two-state, sub-second, idempotent workflow that runs once per transaction. The README's own episode-2 rubric puts this flow on Express.

Evidence. Template as cited; pricing in Appendix C. Live run: ExecutionTime p99 and execution count per hour from the simulator run.

Why it matters. Standard bills USD 25 per million state transitions. At 10 TPS this workflow makes about 78 million transitions a month (three per execution, including start and end), or USD 1,944. Express bills per request and per GB-second: about USD 26 plus USD 3 in duration for the same volume. It is the largest line on the modeled bill by more than an order of magnitude; nothing else in this stack costs more than USD 100 a month.

Recommendation. StateMachineType: EXPRESS, LoggingConfiguration at ERROR level (not ALL, see F-19), and F-03 fixed first so that Express's at-least-once semantics do not hide a publish failure. Idempotency is already guaranteed by the PutItem key (§7), so the switch is safe.

Savings / cost of fix. USD ~1,910/month at 10 TPS (High confidence on the arithmetic, Medium on the assumed rate). A one-line change plus a redeploy.

F-11 — publish-events Lambda is glue that a direct integration replaces

Severity: S2 Area: §3, §4 Effort: S
Impact: Cost (small), latency, and one fewer runtime to keep patched.

What I found. All three workflows invoke publish-events (infrastructure/template.yaml:796-808) solely to call PutEvents. The arn:aws:states:::events:putEvents optimized integration has existed since July 2020, the month this code was last touched.

Recommendation. Replace the three lambda:invoke tasks with events:putEvents tasks and grant events:PutEvents on the bus to each workflow role. Delete the function. This also resolves F-03 structurally.

Savings / cost of fix. At 10 TPS: ~26M invocations/month ≈ USD 5 in requests plus ~USD 9 in duration. Small in money; large in reliability because the failure mode in F-03 disappears.

F-19 — Express workflow logs ALL with execution data, including account numbers

Severity: S2 Area: §3, §8, §9 Effort: S
Impact: Cost Live run: log ingestion GB/month and a compliance exposure.

What I found. customer/template.yaml:574-579: Level: ALL, IncludeExecutionData: true. The event detail for this workflow carries whatever the producer sent; on the transaction path that is from-account and to-account IBANs and amounts. Retention is 7 days (customer/template.yaml:638), which limits the window but not the ingestion cost.

Recommendation. Level: ERROR in production, IncludeExecutionData: false unless a specific debugging window needs it. If execution data must be logged, add a data-protection policy on the log group masking account identifiers.

F-20 — Standard workflows have no logging or tracing

Severity: S3 Area: §3, §8 Effort: S

What I found. TransactionProcess and NormalizeAccountProcess set no LoggingConfiguration and no TracingConfiguration. Execution history is available in the console for 90 days, but nothing reaches CloudWatch Logs or X-Ray, so there is no way to alarm or to correlate with the Lambda that publishes.

Recommendation. Level: ERROR logging to a group with 30-day retention; TracingConfiguration.Enabled: true on both, and Tracing: Active on the three functions.

4.4 Lambda

F-01 — All functions on the deprecated nodejs12.x runtime

Severity: S1 Area: §4, §11 Effort: M
Impact: Operability: the stack cannot be redeployed; security: no runtime patches since 2022.

What I found. infrastructure/template.yaml:801 and simulator/template.yaml:1057 declare Runtime: nodejs12.x for the two functions that have templates. The third function, simulate-transaction, has handler code and a package.json but no template at all; it was deployed by hand, outside IaC (see §4.11). Node 12 reached end of support in Lambda on 2022-03-31; function creation was blocked from 2022-04-30 and updates from 2023-02-28.

Why it matters. The first sam deploy that touches a function resource fails. That is every deploy after any change to these stacks. There is no path to fix F-03 without first fixing this.

Recommendation. Runtime: nodejs22.x and the SDK migration in F-02. Add a CI check that fails on any runtime within six months of its deprecation date.

F-02 — Handlers depend on aws-sdk v2, which current runtimes do not bundle

Severity: S1 Area: §4 Effort: M
Impact: Blocks F-01.

What I found. All three handlers require('aws-sdk') (infrastructure/publish-events/app.js:734, simulator/simulate-approval/app.js:899, simulator/simulate-transaction/app.js:941) and none declares it as a dependency, relying on the runtime's bundled copy. Node 18 and later bundle only SDK v3. aws-sdk v2 entered maintenance mode in September 2024 and end of support in September 2025.

Recommendation. Migrate to @aws-sdk/client-eventbridge and @aws-sdk/client-sfn, declare them in package.json, and bundle with esbuild via SAM's Metadata.BuildMethod: esbuild so the package stays small. The simulate-transaction function also pins faker ^4.1.0, an unmaintained package whose successor is @faker-js/faker.

F-16 — No log retention on Lambda log groups

Severity: S3 Area: §4, §8, §10 Effort: S
Impact: Cost: log storage grows forever. Live run: StoredBytes per group.

What I found. No AWS::Logs::LogGroup is declared for any of the three functions, so Lambda creates them on first invoke with Never expire. publish-events also console.logs every full event payload (app.js:752), which at 10 TPS is ~26M log lines a month carrying account identifiers.

Recommendation. Declare each group with RetentionInDays: 30, log at INFO with structured JSON and no payload bodies, and use the same data-protection policy as F-19.

F-21 — 3-second timeouts on functions that make network calls

Severity: S3 Area: §4 Effort: S

What I found. simulator/template.yaml:1047-1049 sets a global 3-second timeout; publish-events uses SAM's 3-second default. A cold start of the v2 SDK plus one API call fits, usually. Under a Step Functions or SQS retry it becomes a source of Task timed out errors that F-09 does not catch. Live run: Duration p99 and any timeouts.

Recommendation. 10 seconds on publish-events and simulate-approval; keep the client-side SDK timeout shorter than the function timeout.

I-03 — The simulator double-sends and drops events

Severity: Info Area: §4

simulator/simulate-transaction/app.js:960-968 starts a new batch every tenth entry but pushes the batch reference at entry 0 and again at entry 9, so the first ten entries of every second are sent twice; entries at positions 19, 29, … are never pushed. It is a load generator, so this is not a production defect, but any cost or throughput baseline measured with it is off by roughly +10% / −10%. The live figures in this report are taken with it corrected.

4.5 ECS / Fargate

Reviewed; no ECS resources in scope. No findings.

4.6 Data

F-14 — Raw and processed S3 buckets have no lifecycle policy

Severity: S2 Area: §6, §10 Effort: S
Impact: Cost: unbounded growth. Live run: bucket size and object count after the run; projected at 10 TPS ≈ 1.5 GB/month raw and ~43K objects/month (Appendix C).

What I found. billing/template.yaml:71-76: both buckets declared with no properties. The Firehose backup writes every raw record gzipped to RawDataBucket in 1 MB / 60-second buffers, and the Parquet output to ProcessedDataBucket. Nothing expires or tiers.

Recommendation. Lifecycle rule on the raw bucket: transition to Glacier Instant Retrieval at 30 days, expire at 365 (or whatever the audit-retention requirement is). Intelligent-Tiering on the processed bucket. S3BackupMode is worth keeping only if the raw copy has a consumer; otherwise disable it and halve the writes.

F-15 — DynamoDB table with no PITR, no TTL, on-demand by default, and a type mismatch with the Glue schema

Severity: S2 Area: §6 Effort: S
Impact: Reliability (no recovery path); data quality across OLTP/OLAP.

What I found. billing/template.yaml:187-188 uses AWS::Serverless::SimpleTable with no properties: on-demand billing, no PointInTimeRecoverySpecification, no TTL, primary key id as string. The workflow writes total_amount as a DynamoDB string ("S.$", billing/template.yaml:218), while the Glue table declares it double (billing/template.yaml:122-123). The transaction id is the EventBridge event id, not a business identifier (billing/template.yaml:49).

Recommendation. PITR on; TTL if transactions have a retention limit; N for total_amount; a real transaction id in the event detail. On-demand is right for the simulator's bursty pattern; at a steady 10 TPS provisioned capacity with autoscaling is ~USD 7/month vs ~USD 32 on-demand (Appendix C). Only worth doing once the real traffic shape is known.

I-01 — Bucket encryption and public-access block not declared

Severity: Info Area: §6, §9

Neither bucket sets BucketEncryption or PublicAccessBlockConfiguration. In 2020 that was a finding. Since 2023 S3 encrypts new objects with SSE-S3 and blocks public access on new buckets by default, so as deployed today these are safe. Declare both anyway so a policy check can prove it.

4.7 API edge

Reviewed; no API Gateway or Cognito in scope. Producers call PutEvents with IAM credentials. No findings.

4.8 Observability

F-17 — No alarms

Severity: S1 Area: §8 Effort: M
Impact: Every failure mode in this report is silent.

What I found. No AWS::CloudWatch::Alarm in any stack. Nothing pages on failed executions, DLQ depth (there are no DLQs), Lambda errors, Firehose delivery failures, or throttles.

Recommendation. The minimum set, one alarm each: ExecutionsFailed and ExecutionsTimedOut per state machine; ApproximateNumberOfMessagesVisible on every DLQ (after F-06/F-13); Errors and Throttles on each function; DeliveryToS3.Success < 1 on the Firehose; FailedInvocations per rule. Route to an SNS topic that reaches a human. Add a composite alarm for "transaction pipeline unhealthy" so on-call gets one page, not seven.

F-18 — No structured logging or correlation id across the event chain

Severity: S3 Area: §8 Effort: M

What I found. Handlers log with console.log of raw objects. No correlation id is carried from transaction-initiated through the workflow to transaction-processed; the EventBridge event id is available but not propagated into logs or the published event's detail. No X-Ray.

Recommendation. Lambda Powertools logger with the event id as a correlation key; propagate it in every published event's detail; enable tracing (F-20).

4.9 Security and IAM

F-04 — LegacyOperationQueue accepts SendMessage from any AWS principal

Severity: S1 Area: §9 Effort: S
Impact: Security: unauthenticated-by-design message injection into the legacy operations path.

What I found. operations/template.yaml:865-883: the queue policy allows sqs:SendMessage to Principal: {AWS: "*"} with the condition ForAllValues:ArnEquals: {aws:SourceArn: <topic ARN>}. aws:SourceArn is a single-valued key, and per the IAM documentation ForAllValues evaluates to true when the key is absent from the request. A direct SendMessage call from any AWS account carries no aws:SourceArn, so the condition passes and the message is accepted. The intent was ArnEquals (no set-operator prefix), which denies when the key is missing.

Evidence. Template lines cited; IAM docs on multivalued condition keys. Live run: aws sqs send-message from a second account succeeds.

Why it matters. Whatever consumes this queue (F-23) trusts that messages came from the topic. Anyone who learns the queue URL, which is a stack output, can inject arbitrary operation-performed payloads.

Recommendation. Replace ForAllValues:ArnEquals with ArnEquals. Add an explicit Deny for aws:SecureTransport: false. Add a policy lint (cfn-guard or checkov) to CI; this pattern is a known rule in both.

F-22 — Hard-coded resource names prevent multiple environments and safe rollback

Severity: S2 Area: §9, §11 Effort: S
Impact: Operability: no staging stack in the same account; replacement updates fail.

What I found. StateMachineName: TransactionProcess, NormalizeAccountProcess, ExpiredSubscriptionProcess; TopicName: LegacyOperationTopic; QueueName: LegacyOperationQueue; LogGroupName: /customer; SSM parameter /EventBusARN; bus name AnyCompany; Glue database app2025. Any second deployment in the account collides, and any change that requires replacement of a named resource fails CloudFormation's update.

Recommendation. Drop explicit names or suffix them with a stage parameter. Keep the bus name as a parameter so consumers can target AnyCompany-staging.

I-04 — IAM is otherwise well scoped

Severity: Info Area: §9

Every role is per-workflow or per-function, every policy names its resource, the Firehose role uses an external-id condition, and the only * resource is on the log-delivery actions that genuinely do not support resource-level permissions and are commented as such (customer/template.yaml:667-670). This is better than most production stacks.

4.10 Cost

F-24 — No tags, budgets, or anomaly detection

Severity: S3 Area: §10 Effort: S

What I found. No Tags on any resource or stack; no AWS::Budgets::Budget; no Cost Anomaly Detection monitor. Cost cannot be attributed per flow. Live run: untagged share is 100% of the workload's spend.

Recommendation. Stack-level tags (workload, stage, owner) via sam deploy --tags, a monthly budget with an 80% alert, and an anomaly monitor on the linked account.

Cost findings F-05, F-11, F-14, F-15, F-16 and F-19 are filed under their service sections; the roll-up is in §5.

4.11 Delivery and IaC

Findings F-01, F-02, F-22 and F-23 cover this area. Additionally:

Live evidence for this report is collected with a read-only evidence script and deployed from a minimally patched copy of the application; both will be published with the final report so the method is reproducible.


5. Cost analysis

Modeled, not billed. Assumption throughout: 10 transactions/s steady (the simulator's default), 30-day month, us-east-1 pricing as of 2026-09. The live run will replace the modeled column.

Spend by service, modeled at 10 TPS

Service Modeled USD/month Driver Live
Step Functions (Standard) 1,944 78M transitions × USD 25/M pending live run
Lambda ~14 26M invokes + duration pending live run
DynamoDB (on-demand) ~32 26M writes × USD 1.25/M pending live run
Firehose ~30 26M records, 5 KB minimum billable each pending live run
S3 < 1, growing ~3 GB/month new, never expired pending live run
CloudWatch Logs pending live run console.log of every payload + Express ALL pending live run
EventBridge ~26 26M custom events × USD 1/M pending live run
SQS, SNS, Glue < 5 pending live run
Total ~2,050 + logs + S3 growth pending live run

Unit economics

Unit Modeled cost Method
Per transaction USD 0.000079 total ÷ 25.9M
Per transaction after F-05, F-11 USD 0.000005

Savings summary

Finding Monthly saving USD Effort Confidence
F-05 Standard → Express ~1,910 S High on arithmetic, Medium on volume
F-11 Remove glue Lambda ~14 S High
F-14 S3 lifecycle small today; unbounded growth stops S Medium
F-19 Express logging level pending live run S
F-16 Lambda log retention + payload logging pending live run S
F-15 DynamoDB provisioned (only at steady load) ~25 S Low until traffic shape is known
Total ~1,950 + log savings

Commitments

Nothing here is steady enough to justify a Savings Plan until the Standard-to-Express change lands; after it, the workload is under USD 150/month and a commitment is not worth the lock-in.


6. Remediation backlog

Import-ready version: backlog.csv.

Order ID Title Sev Effort Depends on Why this position
1 F-04 Fix queue policy condition operator S1 S — One line; closes an open queue
2 F-01 Migrate runtime to nodejs22.x S1 M F-02 Nothing else can deploy until this lands
3 F-02 Migrate to SDK v3, bundle with esbuild S1 M — Prerequisite for F-01
4 F-03 Throw on PutEvents failure S1 S F-01 Stops silent event loss
5 F-11 Replace publish-events with direct integration S2 S F-03 Makes F-03 structural; do together
6 F-09 Retry/Catch on every task S1 M F-01 Safety net
7 F-06 DLQs on all EventBridge targets S1 M — Safety net
8 F-13 SQS redrive policies and timeouts S2 S — Safety net
9 F-17 Alarms S1 M F-06, F-13 Needs DLQs to alarm on
10 F-05 TransactionProcess → Express S1 S F-03, F-09 Largest saving; safe once publisher fails loudly
11 F-07 Fix Message.? typo S2 S — Real bug, invisible today
12 F-12 Fix batch handling in simulate-approval S2 S F-13
13 F-10 Callback timeout + idempotent execution names S2 S —
14 F-19 Express logging to ERROR, no execution data S2 S — Cost and PII
15 F-14 S3 lifecycle rules S2 S —
16 F-15 DynamoDB PITR, TTL, numeric amount S2 S —
17 F-22 Remove hard-coded names, add stage parameter S2 S — Enables a staging stack for everything above
18 F-16 Lambda log retention, structured logs S3 S F-01
19 F-20 Logging and tracing on Standard workflows S3 S —
20 F-18 Correlation id end to end S3 M F-16
21 F-21 Function timeouts S3 S —
22 F-08 Bus archive and schema registry S3 S —
23 F-23 Bring legacy consumer into IaC or delete S3 S —
24 F-24 Tags, budget, anomaly detection S3 S —

Suggested sequencing


7. What is working


8. Next steps

  1. Q&A call: [date]. Bring whoever owns the customer and billing stacks.
  2. Async follow-up window closes [date + 30 days].
  3. If you want help executing this: the remediation package covers the runtime migration, the publisher fix, retries, DLQs and alarms (F-01, F-02, F-03, F-06, F-09, F-17; the top three findings and their prerequisites) in two weeks for USD 8,000 fixed. Beyond that, USD 1,500/day with a five-day minimum. Reply to this report to start.

Appendix A — Checklist coverage

Live run: full checklist with Done / N/A / Not reviewed per item. Static pass: §0 partial (no account), §1–§4, §6, §9, §11 done from IaC; §5, §7 N/A; §8, §10 partial pending metrics.

Appendix B — Evidence index

Finding Evidence Collected
F-01, F-02 infrastructure/template.yaml:801, simulator/template.yaml:1057, handler require lines 2026-09-12
F-03 infrastructure/publish-events/app.js:755-758 2026-09-12
F-04 operations/template.yaml:865-883; IAM docs, "Multivalued context keys" 2026-09-12
F-05 billing/template.yaml:193-199; Appendix C 2026-09-12
F-06 all four AWS::Events::Rule resources 2026-09-12
F-07 customer/template.yaml:483 2026-09-12
F-09 three DefinitionString blocks 2026-09-12
F-10 customer/template.yaml:477-489 2026-09-12
F-12 simulator/simulate-approval/app.js:901-925; simulator/template.yaml:1058-1062 2026-09-12
F-13 customer/template.yaml:544-546; operations/template.yaml:860-863 2026-09-12
F-14, F-15, I-01 billing/template.yaml:71-76, 187-188, 218, 122-123 2026-09-12
F-19, F-20 customer/template.yaml:574-579; absence in billing and NormalizeAccountWorkflow 2026-09-12
F-22 grep for Name: across templates 2026-09-12
Others pending live run

Line numbers refer to the concatenated template and handler dump: templates.txt.

Appendix C — Cost calculations

F-05, Standard vs Express at 10 TPS. Executions/month = 10 × 86,400 × 30 = 25,920,000. Standard: 3 transitions per execution (start, task 1 → task 2, end) = 77,760,000 transitions × USD 25 / 1,000,000 = USD 1,944. (Free tier of 4,000 transitions ignored.) Express: requests 25.92M × USD 1.00 / 1M = USD 25.92. Duration: assume 64 MB and 100 ms per execution = 0.0625 GB × 0.1 s = 0.00625 GB-s × 25.92M = 162,000 GB-s × USD 0.00001667 = USD 2.70. Total USD ~29. Saving ≈ USD 1,915/month. Scales linearly: at 1 TPS it is USD ~190; at 50 TPS, USD ~9,500.

F-11, glue Lambda. 25.92M invocations × USD 0.20 / 1M = USD 5.18. Duration at 128 MB, 80 ms average: 0.128 × 0.08 = 0.01024 GB-s × 25.92M = 265,000 GB-s × USD 0.0000166667 = USD 4.42. Plus ~USD 4 for the Step Functions Lambda invoke overhead on Express. ≈ USD 14.

F-15, DynamoDB. On-demand: 25.92M writes × USD 1.25 / 1M = USD 32.4 (1 KB items). Provisioned: 10 WCU steady with autoscaling headroom to 15 WCU ≈ 15 × USD 0.00065 × 730 h = USD 7.1. Reads not modeled (none in the stack).

F-14, S3 growth. Raw backup: ~200-byte JSON records gzipped ≈ 60 bytes × 25.92M ≈ 1.5 GB/month; Parquet processed similar. Firehose's 1 MB / 60 s backup buffer at this rate produces ~43K objects/month, which is the real cost: PUT requests ≈ USD 0.22, storage ≈ USD 0.07 per month retained, cumulative. Small in dollars; the finding is that it is unbounded and un-tiered, not that it is expensive today. Revised from the earlier 4–6 GB/day estimate once record sizes were checked.

Appendix D — Inventory export

Live run: aws cloudformation list-stack-resources for the five stacks.

Appendix E — Glossary

Omitted; audience is engineers.


Static review performed against the public repository. Live evidence to be collected under a read-only role in a sandbox account and deleted 30 days after the follow-up window closes. This is a sample report; "AnyCompany" is AWS's fictional client and no real customer data is involved.

Book a fit call