RepoLightning AILightning AIpublished Aug 14, 2026seen 3w

Lightning-AI/litracer

Go

Open original ↗

Captured source

source ↗
published Aug 14, 2026seen 3wcaptured 3whttp 200method plain

Lightning-AI/litracer

Description: A tool for converting Lightning AI's LitData logs to Chrome-compatible trace files.

Language: Go

License: Apache-2.0

Stars: 2

Forks: 0

Open issues: 0

Created: 2026-08-14T10:53:00Z

Pushed: 2026-08-14T11:16:06Z

Default branch: main

Fork: no

Archived: no

README:

Litracer

Convert LitData enable_tracer() logs into Chrome / Perfetto traces.

Repo: Lightning-AI/litracer. LitData cookbook: Debug & Profile.

---

Install

From this repo (recommended while the Go module path is still github.com/deependujha/litracer):

git clone https://github.com/Lightning-AI/litracer.git
cd litracer
go build -o litracer .

Via Go (module path; pulls the published module, not necessarily this GitHub default branch):

go install github.com/deependujha/litracer@latest

Requires Go 1.23+ (go.mod language version). CI and the module toolchain use Go 1.26.6. go install github.com/Lightning-AI/litracer@latest will not work until go.mod is renamed to that module path.

---

Usage with LitData

Call enable_tracer() once per process, before creating the DataLoader. Delete an existing litdata_debug.log before re-tracing (the handler appends).

from litdata.debugger import enable_tracer

enable_tracer(level="chunk", log_file="litdata_debug.log")
# level="batch" | "chunk" (default) | "sample" | "debug" | "off"
# categories=["download", "read", "delete"] # optional override

| Level | What you get | | ----- | ------------ | | batch | Epoch + per-batch spans, plus crashes | | chunk | + download, read, delete, decompress, prefetch | | sample | + per-__getitem__ spans (high volume) | | debug | + .cnt lock refcount spans | | off | Disable |

python train.py
litracer --quiet --validate -o litdata_trace.json litdata_debug.log
litracer --quiet --cat download,read,delete -o io.json litdata_debug.log
# open the json in https://ui.perfetto.dev

--quiet (or TERM=dumb) skips the progress UI and prints a one-line summary with per-category durations. --validate checks the JSON parses and event counts match.

Flags

Usage:
litracer [flags]

Flags:
-h, --help help
-o, --output string Path to the output trace file (default "litdata_trace.json")
-q, --quiet Skip the progress UI and print a one-line summary
--validate Parse the output JSON and exit non-zero if it is invalid
--cat string Keep only these categories (download,read,delete,decompress,batch,sample,epoch,lock,crash)
--no-complete Keep separate B/E events instead of compacting to complete (X) spans
-w, --workers int Parse goroutines (default: GOMAXPROCS)
-s, --sink int Deprecated: ignored

--cat and --no-complete apply in both quiet and TUI modes.

---

Log format

Each non-empty line is one Chrome event as semicolon-separated key: value pairs. LitData writes exactly one line per event (no traceback bodies). Timestamps are Chrome microseconds (created * 1e6); values older than 1e12 are treated as Unix seconds and multiplied.

ts:;PID:; TID:; name: download;ph: B;cat: download;cname: rail_load;chunk: 12;worker_rank: 1;...

Required keys: name, ph. Everything else becomes a Perfetto arg (or a well-known field: pid, tid, cat, cname, s, dur).

| ph | Meaning | | ---- | ------- | | B / E | Begin / end of a duration span (same name + pid + tid) | | X | Complete span (already compacted; dur set) | | I | Instant (crashes). Scope defaults to t (thread) | | M | Metadata (process_name / thread_name) written by Litracer |

Lines without ph (Python Traceback …, blank lines, noise) are skipped.

Matched B/E pairs are compacted to complete (ph: X) spans unless --no-complete. That keeps Perfetto files smaller.

Stable event names (LitData current tracer)

Indexes live in args, not in the event name, so Perfetto groups all downloads together.

| name | cat | When | | ------ | ----- | ---- | | download | download | Chunk GET to cache | | prefetch | download | Async prefetch gather | | read | read | Chunk mmap / decode | | delete | delete | Evict a consumed chunk | | decompress | decompress | zstd (or other) inflate | | batch | batch | DataLoader batch | | dataloader / combined | epoch | Epoch span | | sample | sample | Per-item __getitem__ | | lock | lock | .cnt increment/decrement (level="debug") | | crash | crash | Prepare-thread death (ph: I). Traceback is also printed to stderr, not the log file |

Every event also carries dist_world_size, dist_global_rank, dist_num_nodes, worker_world_size, worker_rank.

Older logs

Still convert:

  • Names like download_chunk_12, read_chunk_3, prepare_chunks_thread_crashed_TypeError
  • Unix-second timestamps (ts:1750000000.5)
  • Missing cat: (inferred from name)

---

Interpreting the summary

lines=698 events=360 skipped=0 B=0 E=0 X=349 I=1 unmatched_B=0 unmatched_E=0 crashes=1 processes=9 duration_ms=… download=…ms read=…ms delete=…ms
  • skipped — lines with no ph (tracebacks, blanks).
  • unmatched_B / unmatched_E — duration events with no pair. A leftover open read at worker shutdown used to cause this; current LitData closes the last read span.
  • crashes — instant crash / *crashed* events. Also look at the training process stderr for [litdata] PrepareChunksThread CRASHED.

Open litdata_trace.json in ui.perfetto.dev (preferred) or chrome://tracing. For files > 2GB, see Perfetto large traces.

---

Develop

go test ./...
go build -o litracer .

CI (.github/workflows/test.yml) runs go test ./... and go build with Go 1.26.6 on pull requests and on pushes to main.

Notability

notability 5.0/10

New repo from Lightning AI but low traction