A small library for tracing Lambda logs by request ID
I built lambda_json_logger, a small library that makes AWS Lambda emit one JSON object per log line. It has no dependencies.
Lambda’s default log output is plain text, with the request ID baked into a prefix at the start of the line. To CloudWatch Logs Insights that prefix is just part of a string, so you cannot filter on it without writing a parse. The request payload is not there at all, which means “what request produced this error?” is unanswerable after the fact.
Writing a parse on every investigation adds up. Assembling regular expressions in the middle of an incident is not what you want to be doing.
The library swaps in a logging formatter that puts the request ID, environment name and event at the top level of a JSON object. Insights parses JSON logs automatically, so the fields are queryable as they are.
Here is what this post covers:
- Using it
- Request IDs drift across warm starts
- Don’t attach too much
- Testing
- Where this sits next to the built-in feature
Using it
Install directly from GitHub:
pip install "lambda_json_logger @ git+https://github.com/yosuke318/lambda_json_logger.git"
Grab the logger at the top of the handler:
from lambda_json_logger import getLambdaJsonLoggerInstance
def lambda_handler(event, context):
logger = getLambdaJsonLoggerInstance(context=context, stage="dev", event=event)
logger.info("processing started")
return {"statusCode": 200}
The emitted JSON looks like this:
| Field | Contents |
|---|---|
time | Timestamp |
level | INFO, ERROR, and so on |
message | The log message |
function_name | Name of the emitting function |
module | Name of the emitting module |
aws_request_id | Request ID, taken from context |
stage | Environment name, passed in optionally |
event | Only present when passed in |
exception | Only on logger.exception() or exc_info=True |
Which makes the Insights side look like this:
fields @timestamp, message, event.httpMethod, event.path
| filter aws_request_id = "8f7e6d5c-..."
| filter level = "ERROR"
| sort @timestamp desc
Nested keys work with dot notation, so filters like event.path like /^\/orders/ are available.
Request IDs drift across warm starts
This is the part that took the most care to get right.
Lambda reuses execution environments, so the logger is shared within the process. If the formatter holds onto the context from when it was first created, every subsequent invocation keeps printing the previous request’s ID.
What makes it nasty is that the logs still look correct. The JSON is well-formed and every field is present. Only the contents are wrong — and since you are using those logs to investigate, they lead you to the wrong conclusion.
The handler is created once, but the formatter is rebuilt and swapped on every call:
if not logger.handlers:
logger.addHandler(StreamHandler())
formatter = LambdaJsonFormatter(context=context, stage=stage, event=event)
for handler in logger.handlers:
handler.setFormatter(formatter)
That is why the per-invocation values live on the formatter instance rather than in a closure. The flip side: you have to call it at the top of the handler every time. Skip it and you get the previous request’s ID.
There is also logger.propagate = False. The Lambda runtime prints anything that reaches the root logger itself, so propagating would emit every line twice.
Don’t attach too much
Passing event puts it on every log line. CloudWatch bills on ingested bytes, so a large event combined with dozens of log lines per invocation inflates volume quickly. If that matters, skip event, emit one line at the start, and correlate on aws_request_id afterwards. Omitting it removes the key entirely.
Tracebacks from logger.exception() land in the exception field. Newlines are escaped as \n, so CloudWatch does not split them into separate events. Values that cannot be serialised are stringified by default=str, so a log call never blows up on them.
A logging call that fails because it cannot log is worse than a lossy one, so this errs on the side of swallowing.
Testing
context is supplied by the Lambda runtime, so the tests use a DummyContext carrying only aws_request_id. Standard output is redirected to a buffer, and each emitted line is parsed with json.loads and checked structurally.
Since “one JSON object per line” is the contract, starting the assertions from “it parses” was the natural fit.
Where this sits next to the built-in feature
AWS shipped Advanced Logging Controls for Lambda on 16 November 2023. Set the log format to JSON in the function configuration and the runtime emits structured logs including requestId and level. Log levels change without touching code, and you can choose the destination log group.
If you are starting fresh, check whether that covers you first. Changing one setting always beats adding a library.
This library earns its place when you want arbitrary fields such as stage, or when you are on a setup that does not use that feature. The runtime’s JSON output has a fixed set of keys, so custom fields mean owning a formatter.