Files
2026-08-27 11:12:22 +02:00

15 KiB

MARTe Benchmark Suite

Overview

marte-benchmark-suite is a validation and benchmarking infrastructure developed around the MARTe2 real-time framework.

The project is designed to build a sequence of progressively more advanced test scenarios while preserving known-good baselines from the previous stages.

The suite does not modify MARTe2 itself.

Instead, it provides:

  • MARTe configuration files;
  • launchers;
  • configuration patchers;
  • runtime orchestration;
  • output parsers;
  • structural validators;
  • semantic validators;
  • execution evidence.

The development philosophy is incremental:

each new stage introduces one new capability while preserving the validated behaviour of the previous stages.


Current project status

Current validated development baseline:

S03-H3
STATUS=CLOSED
BASELINE=KNOWN-GOOD

S03-H3 is currently the authoritative baseline on which the next stages of the benchmark suite are intended to be built.

The next planned stages are:

S04 → Repeatability / Multi-Run
S05 → Deterministic Parametric Generator
S06 → MARTe2 Core Scaling

These stages are not part of the current S03-H3 baseline unless explicitly implemented and accepted later.


Project purpose

The suite is intended to answer increasingly complex questions about a MARTe2 application.

Conceptually:

S00
Can MARTe2 execute the basic reference application?
        ↓
S01
Can execution timing information be observed?
        ↓
S02
Can deterministic binary data be recorded?
        ↓
S03
Can the recorded data be checked semantically?
        ↓
S03-H1
Can different execution/debug profiles be supported?
        ↓
S03-H3
Can those profiles be formalized and validated
against an authoritative runtime contract?
        ↓
S04
Does the same valid workload remain correct
when executed repeatedly?
        ↓
S05
Can deterministic test workloads be generated
from parameters?
        ↓
S06
How does MARTe2 Core behave as workload complexity grows?

High-level architecture

The benchmark suite is layered around the installed MARTe2 runtime.

                    +------------------+
                    |      MARTe2      |
                    |   MARTeApp.ex    |
                    +---------+--------+
                              ^
                              |
                     common/marte_runner.sh
                              ^
                              |
                         launcher/
                              ^
                              |
                 scenario / configuration
                              |
                              v
                      MARTe execution
                              |
                              v
                         RUN DIRECTORY
                              |
                 +------------+-------------+
                 |                          |
                 v                          v
             runtime logs              output files
                 |                          |
                 +------------+-------------+
                              |
                              v
                       parser / validator
                              |
                              v
                       PASS / FAIL result

MARTe2 and MARTe2-components remain external dependencies.

The benchmark suite orchestrates and validates them but is intentionally kept separate from their source trees.


Repository structure

Current source structure:

marte-benchmark-suite/
│
├── common/
│   ├── environment.sh
│   ├── marte_runner.sh
│   └── paths.sh
│
├── configurations/
│   └── generated/
│       ├── S00/
│       ├── S01/
│       ├── S02/
│       ├── S03/
│       ├── S03H1/
│       └── S03H3/
│
├── launcher/
│   ├── marte_benchmark_launcher.sh
│   ├── marte_benchmark_launcher_s03h1.sh
│   ├── marte_benchmark_launcher_s03h3.sh
│   └── marte_benchmark_menu.sh
│
├── patchers/
│   └── configuration/
│
├── scenarios/
│   ├── S01_timing_display/
│   ├── S02_binary_recording/
│   ├── S03_semantic_sentinels/
│   ├── S03H1_execution_profiles/
│   └── S03H3_execution_profiles/
│
└── runs/

runs/ contains execution outputs and is intentionally excluded from normal Git version control.

The configurations under configurations/generated/ are NOT generic build artifacts. They represent configuration baselines produced by controlled project stages and are therefore versioned.


Common runtime layer

common/paths.sh

Centralizes the paths used by the benchmark infrastructure.

Its purpose is to avoid hard-coding the same MARTe and project locations in multiple scripts.


common/environment.sh

Defines the common execution environment used by the benchmark scripts.

Keeping environment handling in one place improves reproducibility and avoids different scenarios silently using different runtime assumptions.


common/marte_runner.sh

Provides the common MARTe execution mechanism shared by benchmark stages.

Conceptually:

launcher
   ↓
common runner
   ↓
MARTeApp.ex
   ↓
run directory
   ↓
logs / outputs

Higher-level stages should reuse the authoritative runner rather than create alternative MARTe execution paths unless a future specification explicitly authorizes such a change.


Launchers

The launcher/ directory contains the user-facing execution layer.

The project currently includes launchers from different development stages, including the S03-H1 and S03-H3 generations.

The current S03-H3 architecture supports execution profiles such as:

MINIMAL
TEST
DIAGNOSTIC

The profiles allow the same benchmark infrastructure to expose different levels of runtime/debug functionality without changing the fundamental correctness contract.


Patchers

Configuration patchers live under:

patchers/configuration/

They transform an already-known configuration into the configuration required by the next development stage.

The dependency chain currently visible in the repository is:

S00
 ↓
build_s01_from_s00.sh
 ↓
S01
 ↓
build_s02_from_s01.sh
 ↓
S02
 ↓
build_s03_from_s02.sh
 ↓
S03

Additional controlled transformations support S03-H1 and S03-H3.

This model helps preserve provenance:

new configuration
=
known parent
+
explicit transformation

rather than independently rewriting every configuration from scratch.


Validation layer

The suite does not consider a MARTe process exit by itself sufficient proof that a test succeeded.

Validation is layered.

Depending on the stage, the suite can inspect:

configuration structure
        +
runtime execution
        +
generated files
        +
semantic content
        +
expected sentinels

and only then derive the corresponding result.


S00 — Core reference

S00 provides the basic MARTe2 reference configuration:

configurations/generated/S00/smoke_s0_core.marte

It acts as the root from which later benchmark configurations are progressively derived.


S01 — Timing

S01 introduces timing-related observation on top of the S00 reference configuration.

Relevant current files include:

patchers/configuration/build_s01_from_s00.sh
configurations/generated/S01/from_patcher/
scenarios/S01_timing_display/validate.sh

S01 established infrastructure that later stages could reuse rather than reimplement.


S02 — Binary recording

S02 extends the previous stage with binary data recording.

Relevant components include:

patchers/configuration/build_s02_from_s01.sh

configurations/generated/S02/from_patcher/
    smoke_s2_binary_recording_from_s01.marte

scenarios/S02_binary_recording/validate.sh

Binary recording provides machine-readable data that later stages can parse and validate.


S03 — Semantic validation

S03 adds semantic verification of the generated data.

Relevant components include:

patchers/configuration/build_s03_from_s02.sh

configurations/generated/S03/from_patcher/
    smoke_s3_semantic_sentinels_from_s02.marte

scenarios/S03_semantic_sentinels/
    parse_s03_binary.py
    validate.sh

The key conceptual change is:

"data file exists"

is not enough.

The project also checks whether the data means what the test expected it to mean.


S03-H1 — Execution profiles

S03-H1 introduced explicit execution profiles and supporting validation infrastructure.

Current repository elements include:

launcher/marte_benchmark_launcher_s03h1.sh

patchers/configuration/
    build_s03h1_from_s02.sh

scenarios/S03H1_execution_profiles/
    export_s03_csv.py
    validate.sh
    verify_s03h1_config.py

The historical S03-H1 configurations currently retained are:

TEST
DIAGNOSTIC

They remain important because S03-H3 intentionally reuses validated historical configurations where appropriate instead of duplicating them unnecessarily.


S03-H3 — Current known-good baseline

S03-H3 is the current closed benchmark stage.

Its role is to formalize the execution-profile architecture and protect the runtime contract used by future development.

Relevant repository elements include:

launcher/
    marte_benchmark_launcher_s03h3.sh
    marte_benchmark_menu.sh

patchers/configuration/
    build_s03h3_minimal_from_s03h1_test.sh

configurations/generated/S03H3/from_patcher/
    smoke_s03h3_execution_profiles_minimal_from_s03h1_test.marte

scenarios/S03H3_execution_profiles/
    validate.sh
    verify_s03h3_minimal_structure.py

S03-H3 defines three go-forward execution profiles:

MINIMAL
TEST
DIAGNOSTIC

The MINIMAL configuration is derived from the validated S03-H1 TEST configuration by explicitly removing unnecessary display functionality while preserving the required execution behaviour.

TEST and DIAGNOSTIC reuse their validated historical configurations.

This avoids unnecessary duplication and reduces the number of independent configuration variants that must be trusted.


Why S03-H3 matters

Before advancing toward larger benchmark campaigns, the project needs one well-understood baseline.

S03-H3 provides that reference point.

Conceptually:

KNOWN INPUT
    ↓
KNOWN CONFIGURATION
    ↓
KNOWN LAUNCHER
    ↓
KNOWN RUNTIME DEPENDENCIES
    ↓
KNOWN VALIDATION
    ↓
KNOWN-GOOD RESULT

Future stages should extend this baseline rather than silently redefine it.


Runtime dependencies

The benchmark suite depends on an external MARTe installation.

The current runtime environment includes:

MARTe2
MARTe2-components

Important runtime objects include MARTe executables/libraries and component libraries required by the active configurations.

These external projects are not copied into this repository.

This repository therefore acts as:

benchmark source + configuration + validation infrastructure

and not as a complete binary backup of MARTe.

A separate verified restore mechanism is required for complete disaster recovery of the known-good runtime environment.


Run directories

Runtime execution artifacts are written under:

runs/

A run directory may contain logs, generated data and validation evidence.

These files can be large and are produced repeatedly.

For this reason:

runs/

is intentionally excluded through .gitignore.

Git is used for source/configuration history.

Runtime evidence and full environment recovery are handled separately.


Git versus full restore backup

The Git repository and the full restore point serve different purposes.

Git
│
├── scripts
├── configurations
├── validators
├── parsers
├── documentation
└── development history

while a complete restore point additionally needs the external runtime environment required to reproduce the validated MARTe execution.

Therefore:

Git backup ≠ complete runtime restore point

Both mechanisms are complementary.


Development governance

The project follows a controlled development process.

A typical new stage follows approximately:

DESIGN / SPECIFICATION
        ↓
independent technical review
        ↓
technical arbitration
        ↓
human freeze
        ↓
implementation
        ↓
tests
        ↓
independent implementation review
        ↓
technical arbitration
        ↓
human acceptance
        ↓
new known-good baseline

A development-stage implementation is therefore not automatically considered part of the authoritative baseline simply because code exists.


Planned roadmap

S04 — Repeatability / Multi-Run

Main question:

Does the same already-valid S03-H3 workload remain correct when executed repeatedly?

Conceptually:

S03-H3 run #1 → result
S03-H3 run #2 → result
S03-H3 run #3 → result
...
S03-H3 run #N → result
        ↓
campaign-level repeatability result

S04 is primarily a correctness/repeatability stage.

It is not intended to turn repeatability percentages or CycleTime observations into performance claims.


S05 — Deterministic Parametric Generator

Main question:

Can test workloads be generated automatically from controlled parameters while remaining reproducible?

Conceptually:

parameters
+
seed
   ↓
generator
   ↓
deterministic MARTe scenario

The same inputs should regenerate the same intended scenario.


S06 — MARTe2 Core Scaling

Main question:

How does the MARTe2 Core workload behave as controlled complexity increases?

S06 is expected to reuse the capabilities established by the preceding stages:

S03-H3 → correctness baseline
S04    → repeatable execution
S05    → deterministic workload generation
S06    → controlled scaling study

Current boundary

At the time of this README baseline:

S03-H3 = CLOSED / KNOWN-GOOD

S04 = next development target
S05 = planned
S06 = planned

The existence of this roadmap must not be interpreted as evidence that S04, S05 or S06 have already been implemented or validated.


Repository philosophy

The repository should remain:

small
reviewable
deterministic
traceable
reproducible

Generated runtime outputs should not be mixed with source history.

Known-good configurations, however, may intentionally be versioned even when they were produced by controlled patchers, because they are part of the validated benchmark baseline.


License and distribution

The licensing and distribution policy for this benchmark suite must be defined by the project owners before public redistribution.

MARTe2 and MARTe2-components retain their respective upstream licensing and ownership terms.