Protocol Buffer Configuration
DR_EVT supports structured configuration files using Protocol Buffers (protobuf) for complex simulation setups.
Why Use Protobuf Config?
Benefits over command-line arguments:
✅ Reproducible: Configuration files can be versioned and shared
✅ Complex setups: Manage many parameters in one file
✅ Type-safe: Protobuf validates types and required fields
✅ Documented: Schema defines all available options
✅ Composable: Override config with command-line arguments
Configuration File Format
Configuration files use Protocol Buffer text format (.textproto extension).
Basic Example
sim_config.textproto:
infile: "trace.csv"
outfile: "results.csv"
total_nodes: 1000
backfill_policy: "easy"
priority_policy: "fcfs"
Run with:
${CMAKE_INSTALL_PREFIX}/bin/simulator --config sim_config.textproto trace.csv
Complete Example
advanced_config.textproto:
# Input/Output
infile: "workload.csv"
outfile: "schedule_output.csv"
resource_trace: "node_availability.csv"
# System Configuration
total_nodes: 2000
seed: 42
# Scheduling Policies
backfill_policy: "easy" # Options: "easy", "conservative", "none"
priority_policy: "fcfs" # Options: "fcfs", "sjf", "ljf"
# Queue Implementation (FCFS scheduler only)
queue_impl: "circular" # Options: "circular", "deque", "multimap", "block"
wait_queue_capacity: 0 # 0 = size of job trace; only used when queue_impl="circular"
wait_queue_overflow: "grow" # "abort" | "grow"; only used when queue_impl="circular"
# Job-record store (Trace::m_data, a boost::circular_buffer bounding
# memory via front-only eviction - see
# docs/dev/design-decisions/OUT_TRACE_STREAMING.md for the design)
job_store_capacity: 0 # 0 = size of job trace
job_store_overflow: "grow" # "abort" | "grow"
# Resource-history circular buffer (bounds memory for --resource_trace)
resource_history_capacity: 0 # 0 = 2x loaded jobs, floored at 4096
# Trace Format
trace_format: "simple" # Options: "simple", "lassen"
timestamp_format: "epoch" # Options: "epoch", "iso"
# Simulation Limits
max_jobs: 100000
max_time: 86400 # Stop after 86400 seconds (24 hours)
# Duration Simulation
run_time_mode: "distribution" # Options: "actual", "distribution", "limit"
run_time_distribution: "normal" # Options: "normal", "lognormal", "uniform"
run_time_scale: 0.8 # Jobs run for 80% of time_limit on average
run_time_stddev: 0.1 # Standard deviation: 10%
# Output Options
verbose: false
Run with:
${CMAKE_INSTALL_PREFIX}/bin/simulator workload.csv --config advanced_config.textproto
Note the positional trace-file argument is still required, matching
infile’s own value inside the config - the positional argument always
wins over whatever infile is set to (by -i/--infile or a config
file), so it must be given on the command line even though the config
already sets it. See --infile_list if you’d rather
avoid a positional argument entirely (mutually exclusive with one).
Configuration Options Reference
Input/Output Parameters
Field |
Type |
Description |
|---|---|---|
|
string |
Input trace file path (required, unless |
|
string |
Path to a file listing multiple trace files, one per line - progressive loading, so |
|
string |
Output schedule file path (default: |
|
string |
Node availability trace (optional) |
System Configuration
Field |
Type |
Default |
Description |
|---|---|---|---|
|
int32 |
Required |
Total compute nodes in system |
|
int32 |
Random |
Random seed for reproducibility |
Scheduling Policies
Field |
Type |
Default |
Options |
|---|---|---|---|
|
string |
|
|
|
string |
|
|
|
string |
|
|
|
uint32 |
|
Power of 2; only used when |
|
uint64 |
|
|
|
string |
|
|
|
uint64 |
|
|
|
string |
|
|
|
double |
|
Refuse to grow the job store past this fraction of available memory (Linux only; must be |
|
uint64 |
|
|
backfill_policy:
"easy"- EASY backfilling (only first queued job gets reservation)"conservative"- Conservative backfilling (all queued jobs get reservations)"none"- No backfilling
priority_policy:
"fcfs"- First-Come-First-Served (arrival order)"sjf"- Shortest Job First (by run time estimate)"ljf"- Longest Job First (by run time estimate)
queue_impl (FCFS scheduler only - SJF/LJF always use multimap):
"circular"- boost::circular_buffer-based (default; measured 14-28% faster thandeque- see../dev/design-decisions/CIRCULAR_QUEUE.md)"deque"- std::deque-based (simple, well-tested fallback)"multimap"- std::multimap-based (for differential testing)"block"- block-based with multi-index (reference implementation, not recommended for performance - see../dev/design-decisions/BLOCK_QUEUE.md)
Trace Format
Field |
Type |
Default |
Options |
|---|---|---|---|
|
string |
|
|
|
string |
|
|
|
string |
|
Any IANA timezone (e.g., |
trace_format:
"simple"- 7-column CSV format (job_submit_time, begin_time, end_time, num_nodes, exit_status, queue, time_limit)"lassen"- 33-column LLNL HPC trace format
timestamp_format:
"epoch"- Integer seconds since Unix epoch (1970-01-01)"iso"- ISO 8601 format (e.g.,"2024-01-15T08:00:00")
Simulation Limits
Field |
Type |
Default |
Description |
|---|---|---|---|
|
int32 |
Unlimited |
Stop after processing N jobs |
|
double |
Unlimited |
Stop after N seconds simulation time |
Run Time Simulation
Control how a job’s actual, observed execution length is determined in simulation mode:
Field |
Type |
Default |
Description |
|---|---|---|---|
|
string |
|
How to determine the job’s actual run time |
|
string |
|
Statistical distribution for sampling |
|
double |
|
Scale factor for run times |
|
double |
|
Standard deviation (for distributions) |
run_time_mode:
"actual"- Read job’s actual run time from trace’sactual_run_timecolumn (default, most realistic)"distribution"- Sample from statistical distribution aroundtime_limit × run_time_scale"limit"- Jobs run for exactlytime_limit(debugging only, unrealistic)
run_time_distribution:
"normal"- Normal distribution: mean=time_limit × run_time_scale, stddev=run_time_stddev"lognormal"- Log-normal distribution with median=time_limit × run_time_scale"uniform"- Uniform distribution: [0,time_limit × run_time_scale]
Example: Realistic Run Time Variation
run_time_mode: "distribution"
run_time_distribution: "normal"
run_time_scale: 0.8 # Jobs run for 80% of time_limit on average
run_time_stddev: 0.1 # ±10% variation
Output Options
Field |
Type |
Default |
Description |
|---|---|---|---|
|
bool |
|
Print detailed scheduling events |
Command-Line Override
Command-line arguments override protobuf config values:
# Config file says total_nodes: 1000
${CMAKE_INSTALL_PREFIX}/bin/simulator trace.csv --config sim_config.textproto --total_nodes 2000
# Result: Uses 2000 nodes (command-line wins)
Precedence (highest to lowest):
Command-line arguments
Protobuf config file (
--config)Built-in defaults
Common Configurations
Production Replay
Replay exactly what happened on a real system:
replay.textproto:
infile: "production_trace.csv"
outfile: "results.csv"
total_nodes: 2048
run_time_mode: "actual" # Use actual run times from trace
backfill_policy: "easy"
priority_policy: "fcfs"
trace_format: "lassen"
timestamp_format: "iso"
timezone: "America/Los_Angeles"
What-If Analysis
Simulate how system would behave with different policy:
what_if.textproto:
infile: "production_trace.csv"
outfile: "what_if_results.csv"
total_nodes: 2048
# Simulation mode with realistic variation
run_time_mode: "distribution"
run_time_distribution: "normal"
run_time_scale: 0.85
run_time_stddev: 0.15
# Try conservative backfilling instead of EASY
backfill_policy: "conservative"
priority_policy: "fcfs"
trace_format: "simple"
timestamp_format: "epoch"
Capacity Planning
Test if system can handle increased load:
capacity_test.textproto:
infile: "synthetic_high_load.csv"
outfile: "capacity_results.csv"
# Test with fewer nodes
total_nodes: 1500
run_time_mode: "limit"
backfill_policy: "easy"
priority_policy: "fcfs"
# Stop after 7 days simulation time
max_time: 604800
verbose: true
Performance Testing
Benchmark different queue implementations:
circular_queue_test.textproto (default, typically fastest):
infile: "large_scale_10k_jobs.csv"
outfile: "circular_queue_results.csv"
total_nodes: 1000
backfill_policy: "easy"
priority_policy: "fcfs"
# queue_impl defaults to "circular" - explicit here for clarity.
# wait_queue_capacity/wait_queue_overflow are optional; omitting them
# defaults to a capacity sized to the job trace, which can never
# overflow.
queue_impl: "circular"
run_time_mode: "limit"
trace_format: "simple"
timestamp_format: "epoch"
block_queue_test.textproto (reference implementation, not recommended
for performance):
infile: "large_scale_10k_jobs.csv"
outfile: "block_queue_results.csv"
total_nodes: 1000
backfill_policy: "easy"
priority_policy: "fcfs"
# Use block queue with a 128-job block size
queue_impl: "block"
block_size: 128
run_time_mode: "limit"
trace_format: "simple"
timestamp_format: "epoch"
Protocol Buffer Schema
The full schema is defined in src/proto/dr_evt_params.proto:
message Simulation_Params {
// Random seed (default: system clock-dependent if unset)
uint32 seed = 1;
// Simulation limits
uint32 max_jobs = 2;
double max_time = 3;
// Input/Output
string infile = 4;
// Path to a file listing multiple trace files, one per line -
// progressive loading (see docs/dev/design-decisions/
// OUT_TRACE_STREAMING.md): each is loaded in turn as the simulation
// reaches it, so job_store_capacity can actually bound memory.
// Mutually exclusive with infile - do not set both.
string infile_list = 26;
string outfile = 5;
string resource_trace = 6;
// Enable verbose output for debugging/testing (default: false)
bool verbose = 7;
// Scheduling parameters
int32 total_nodes = 8; // default: 795
string backfill_policy = 9; // "easy", "conservative", or "none" (default: "easy")
string priority_policy = 10; // "fcfs", "sjf", or "ljf" (default: "fcfs")
// Trace format
string trace_format = 11; // "simple" or "lassen" (default: "simple")
string timestamp_format = 12; // "epoch" or "iso" (default: "iso")
string timezone = 13; // e.g. "UTC", "America/Los_Angeles"
// Duration simulation
string run_time_mode = 14; // "actual", "distribution", or "limit" (default: "actual")
string run_time_distribution = 15; // "normal", "lognormal", or "uniform" (default: "normal")
double run_time_scale = 16; // default: 1.0
double run_time_stddev = 17; // default: 0.0
// Queue implementation (FCFS scheduler only)
string queue_impl = 18; // "circular", "deque", "multimap", or "block" (default: "circular")
uint32 block_size = 19; // power of 2 (default: 128); only used when queue_impl="block"
uint64 wait_queue_capacity = 20; // 0 = size of job trace (default: 0); only used when queue_impl="circular"
string wait_queue_overflow = 21; // "abort" or "grow" (default: "grow"); only used when queue_impl="circular"
// Job-record store (field 22, formerly "job_store" - a vector/circular
// runtime choice - is reserved, not reused: Trace::m_data is now
// unconditionally boost::circular_buffer, with no vector path to
// choose between)
uint64 job_store_capacity = 23; // 0 = size of job trace (default: 0)
string job_store_overflow = 24; // "abort" or "grow" (default: "grow")
// Refuse to grow the job store past this fraction of available
// memory (Linux only; a no-op elsewhere). Must be > 0.0 and <= 1.0
// (e.g. 0.8); 0.0 (default) disables the check. Independent of
// job_store_overflow.
double memory_pressure_fraction = 27; // default: 0.0 (disabled)
// Resource-history circular buffer (bounds memory for --resource_trace)
uint64 resource_history_capacity = 25; // 0 = 2x loaded jobs, floored at 4096 (default: 0)
}
This mirrors src/proto/dr_evt_params.proto’s actual Simulation_Params message -
check that file directly if this drifts out of sync again.
Validation
Protobuf validates:
Type checking:
total_nodesmust be integer, not stringRequired fields: Missing required fields cause errors
Enum values: Invalid policy names are rejected
Example error:
Error parsing config file: Unknown field "totalnodes" (did you mean "total_nodes"?)
See Also
CLI Options - Command-line alternatives to protobuf config
User Guide Overview - Trace formats and simulation modes
Quick Start - Basic usage examples