Jump to content

Performance

From RidgeRun Developer Wiki

🚧 Documentation is under development

The Video Stitching for Embedded Systems guide is currently under active development. Some sections may be incomplete or change without notice.

Questions? Contact RidgeRun or email to support@ridgerun.com.





Performance

This page contains the performance measurements collected for the Stitcher on the supported embedded platforms.

The main benchmarks use synthetic RGBA inputs so the measurements reflect stitching throughput without adding decoding, encoding, display, or file I/O to the measured path.

Benchmark Methodology

Test Workload

The synthetic benchmarks use videotestsrc to feed RGBA frames directly into the Stitcher. The output is discarded with fakesink sync=false.

The tested input resolutions are:

  • 1280x720
  • 1920x1080
  • 3840x2160

The tested input counts are:

  • 2 cameras
  • 3 cameras
  • 4 cameras
  • 6 cameras

All inputs use RGBA and 30/1 framerate caps.

The sources are non-live and the output sink is not synchronized, so the benchmark is not limited to 30 FPS. Values above 30 FPS mean the pipeline can process frames faster than real time for that configuration.

The system-memory path uses rrstitcher:

videotestsrc ! \
    video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1 ! \
    queue ! stitcher.sink_N

The GL-memory path uses rrglstitcher:

videotestsrc ! \
    video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1 ! \
    queue ! glupload ! \
    video/x-raw(memory:GLMemory),format=RGBA ! \
    queue ! stitcher.sink_N

The glupload stage is required to provide GstGLMemory to rrglstitcher.

Calibration and Camera Layout

The same calibration geometry was used on all platforms for a given resolution and camera count.

The three-camera 1920x1080 calibration is the reference case. The 720p and 4K calibrations use the same geometry with the homographies scaled to the corresponding input resolution.

The two-camera case uses cameras 0 and 1 from the reference calibration.

The four- and six-camera benchmarks use synthetic extensions of the three-camera calibration. Additional camera positions are created by translating the last homography. These cases measure scaling with additional inputs and larger panoramas; they are not physical four- or six-camera calibrations.

The resulting panorama dimensions are:

Inputs 1280x720 1920x1080 3840x2160
2 1693x959 2539x1438 5078x2875
3 2270x1145 3406x1718 6812x3436
4 2849x1145 4274x1718 8547x3436
6 4005x1145 6008x1718 12015x3436

Overlap is defined by the calibration. A fixed overlap percentage was not used as part of this benchmark.

Measurement Method

Each configuration was warmed up once before collecting results. Recorded configurations were then run at least three times sequentially.

A run was only accepted if it processed all requested frames and reached EOS without an error.

FPS was obtained from the final mean_fps reported by the GStreamer perf element where available.

CPU and RAM values are whole-system measurements, not Stitcher process usage.

For Jetson AGX Orin, CPU, RAM, and GPU usage were collected with tegrastats at 500 ms intervals.

For DragonWing and i.MX95, CPU and RAM were sampled from /proc/stat and /proc/meminfo every 500 ms.

DragonWing GPU utilization was obtained from the KGSL device-wide busy counter.

A reliable GPU utilization counter was not available on i.MX95, so GPU usage is marked as unavailable.

The test system was kept otherwise idle while collecting each result.

Recorded-Video Measurement

A smaller recorded-video benchmark was performed to compare the synthetic three-input 1920x1080 results against pipelines using actual camera recordings.

The same three recordings and three-camera calibration are used throughout the comparison. The files are read with filesrc and the output is discarded with fakesink sync=false, so the pipeline runs as fast as possible rather than at the normal playback rate of the recordings.

The test does not include output encoding, display, muxing, or file output.

The input processing path depends on the platform:

Platform Decode and conversion path
NVIDIA Jetson AGX Orin Hardware H.264 decoding with nvv4l2decoder and conversion to RGBA with nvvidconv.
Qualcomm DragonWing Software H.264 decoding with avdec_h264 and videoconvert.
NXP i.MX95 Hardware H.264 decoding with v4l2h264dec, followed by glupload and glcolorconvert. rrstitcher additionally uses gldownload before its system-memory input.

The recorded-video tests are intended as a practical comparison with the synthetic workload. They are not a separate camera-count or resolution performance matrix.

Latency Measurement Method

Latency was measured separately from the throughput benchmark using 1920x1080 RGBA synthetic inputs with 30/1 caps. The latency benchmark covers 2, 3, and 6 inputs.

For rrstitcher, an identity marker is placed immediately before each Stitcher sink. For rrglstitcher, the marker is placed after glupload, so the GL upload is not included in the reported latency.

A probe at each input marker and a probe at stitcher.src record a monotonic timestamp when each buffer passes. Input and output buffers are matched using their PTS.

For each output frame, latency is calculated as the largest input-to-output delay across all camera inputs:

Latency = max(output probe time - input marker probe time)

Tested Platform Configurations

The benchmark used the following platform configurations:

Platform Stitcher configuration Test conditions
NVIDIA Jetson AGX Orin rrstitcher with separate CUDA and OpenGL LibPanorama builds; rrglstitcher with OpenGL MAXN power mode. Tested with dynamic clocks and with jetson_clocks.
Qualcomm DragonWing rrstitcher with OpenGL; rrglstitcher with OpenGL CPU/RAM sampled from /proc. GPU sampled through KGSL.
NXP i.MX95 rrstitcher with OpenGL; rrglstitcher with OpenGL Active Weston session. CPU/RAM sampled from /proc. GPU usage unavailable.

For the Orin jetson_clocks runs, the CPU cores operated at approximately 2201.6 MHz, the GPU at 1300.5 MHz, and EMC at 3199 MHz.

The Orin latency measurements were collected in MAXN mode with jetson_clocks enabled.

Results by Platform and Configuration

FPS values are rounded to the nearest whole number in the main synthetic throughput tables below.

30 FPS Capability Summary

The table below shows the highest tested camera count that reached at least 30 FPS for each input resolution.

Platform Element / Backend 1280x720 1920x1080 3840x2160
DragonWing rrstitcher / OpenGL 6* 4 —
rrglstitcher / OpenGL 6* 4 2
i.MX95 rrstitcher / OpenGL 2 — —
rrglstitcher / OpenGL 4 3 —
Orin AGX MAXN rrstitcher / CUDA 6* 6* 2
rrstitcher / OpenGL 6* 6* 2
rrglstitcher / OpenGL 6* 6* 3
Orin AGX MAXN + jetson_clocks rrstitcher / CUDA 6* 6* 3
rrstitcher / OpenGL 6* 6* 3
rrglstitcher / OpenGL 6* 6* 3

* Six cameras was the highest input count tested. These entries should not be interpreted as a six-camera product limit.

A dash means that the two-camera case did not reach 30 FPS at that resolution.

This table can also be used to estimate the highest tested resolution for a required camera count. For example, the DragonWing rrglstitcher reached 30 FPS with two 4K inputs, while four inputs reached 30 FPS up to 1080p.

Qualcomm DragonWing

Resolution Inputs rrstitcher rrglstitcher
1280x720 2 131 FPS 109 FPS
3 104 FPS 80 FPS
4 90 FPS 74 FPS
6 61 FPS 50 FPS
1920x1080 2 78 FPS 96 FPS
3 49 FPS 57 FPS
4 37 FPS 45 FPS
6 22 FPS 27 FPS
3840x2160 2 27 FPS 41 FPS
3 15 FPS 24 FPS
4 13 FPS 20 FPS
6 7 FPS 12 FPS

NXP i.MX95

Resolution Inputs rrstitcher rrglstitcher
1280x720 2 52 FPS 85 FPS
3 29 FPS 42 FPS
4 20 FPS 36 FPS
6 14 FPS 21 FPS
1920x1080 2 25 FPS 69 FPS
3 15 FPS 34 FPS
4 9 FPS 22 FPS
6 7 FPS 11 FPS
3840x2160 2 7 FPS 20 FPS
3 4 FPS 10 FPS
4 3 FPS Not measurable
6 Unsupported (OOM) Not tested

NVIDIA Jetson AGX Orin

MAXN - Dynamic Clocks

Resolution Inputs rrstitcher CUDA rrstitcher OpenGL rrglstitcher
1280x720 2 211 FPS 278 FPS 294 FPS
3 150 FPS 157 FPS 149 FPS
4 115 FPS 112 FPS 121 FPS
6 77 FPS 71 FPS 72 FPS
1920x1080 2 109 FPS 133 FPS 171 FPS
3 68 FPS 85 FPS 99 FPS
4 55 FPS 57 FPS 80 FPS
6 36 FPS 37 FPS 58 FPS
3840x2160 2 32 FPS 41 FPS 51 FPS
3 21 FPS 28 FPS 48 FPS
4 18 FPS 21 FPS 21 FPS
6 13 FPS 14 FPS 14 FPS

MAXN with jetson_clocks

Resolution Inputs rrstitcher CUDA rrstitcher OpenGL rrglstitcher
1280x720 2 437 FPS 401 FPS 618 FPS
3 280 FPS 262 FPS 367 FPS
4 213 FPS 199 FPS 311 FPS
6 139 FPS 126 FPS 180 FPS
1920x1080 2 200 FPS 197 FPS 248 FPS
3 128 FPS 125 FPS 205 FPS
4 96 FPS 93 FPS 120 FPS
6 62 FPS 56 FPS 86 FPS
3840x2160 2 53 FPS 54 FPS 64 FPS
3 33 FPS 33 FPS 57 FPS
4 25 FPS 24 FPS 28 FPS
6 16 FPS 15 FPS 20 FPS

Synthetic vs Recorded Video

The table below compares the synthetic three-input 1920x1080 workload with the recorded-video pipelines.

The synthetic result isolates the Stitcher, while the recorded-video result also includes the platform's decode and input-conversion path.

Platform / Configuration Element / Backend Decode Synthetic Recorded video
DragonWing rrstitcher / OpenGL Software 49 FPS 52 FPS
rrglstitcher / OpenGL Software 57 FPS 63 FPS
i.MX95 rrstitcher / OpenGL Hardware 15 FPS 16 FPS
rrglstitcher / OpenGL Hardware 34 FPS 26 FPS
Orin AGX MAXN rrstitcher / CUDA Hardware 68 FPS 88 FPS
rrstitcher / OpenGL Hardware 85 FPS 93 FPS
rrglstitcher / OpenGL Hardware 99 FPS 117 FPS
Orin AGX MAXN + jetson_clocks rrstitcher / CUDA Hardware 128 FPS 123 FPS
rrstitcher / OpenGL Hardware 125 FPS 122 FPS
rrglstitcher / OpenGL Hardware 205 FPS 124 FPS

Resource Usage - Three 1920x1080 Synthetic Inputs

This table shows resource usage for the three-input 1080p synthetic benchmark.

These measurements were collected while the pipelines were running at maximum throughput. They do not represent a pipeline rate-limited to 30 FPS.

Platform Element / Backend FPS CPU Avg. RAM Avg. / Peak GPU Avg. / Peak
DragonWing rrstitcher / OpenGL 49 14.50% 5.06% / 5.44% 43.02% / 48.68%
DragonWing rrglstitcher / OpenGL 57 13.55% 5.52% / 6.36% 47.91% / 56.72%
i.MX95 rrstitcher / OpenGL 15 16.01% 18.49% / 20.07% Unavailable
i.MX95 rrglstitcher / OpenGL 34 18.51% 17.21% / 21.60% Unavailable
Orin AGX MAXN rrstitcher / CUDA 68 8.64% 6.40% / 6.52% 55.14% / 88%
Orin AGX MAXN rrstitcher / OpenGL 85 9.54% 9.41% / 9.47% 54.00% / 84%
Orin AGX MAXN rrglstitcher / OpenGL 99 15.13% 6.17% / 6.58% 47.60% / 89%
Orin AGX MAXN + jetson_clocks rrstitcher / CUDA 128 9.88% 8.55% / 9.28% 28.14% / 50%
Orin AGX MAXN + jetson_clocks rrstitcher / OpenGL 125 11.34% 9.41% / 9.51% 34.53% / 59%
Orin AGX MAXN + jetson_clocks rrglstitcher / OpenGL 205 15.73% 8.32% / 9.07% 43.17% / 78%

CPU and RAM values include the rest of the system. GPU values are device-wide and may include work outside the Stitcher.

Latency

The following table shows the measured Stitcher input-to-output latency for 1920x1080 RGBA synthetic inputs.

Values are shown as median / p95 in milliseconds.

Platform Element / Backend 2 Inputs 3 Inputs 6 Inputs
DragonWing rrstitcher / OpenGL 33.340 / 45.404 ms 54.601 / 72.864 ms 122.425 / 165.895 ms
rrglstitcher / OpenGL 23.353 / 28.123 ms 41.104 / 50.244 ms 101.083 / 115.283 ms
i.MX95 rrstitcher / OpenGL 125.350 / 138.248 ms 221.081 / 229.312 ms 450.882 / 454.592 ms
rrglstitcher / OpenGL 41.425 / 43.229 ms 85.984 / 89.472 ms 272.884 / 278.165 ms
Orin AGX MAXN + jetson_clocks rrstitcher / CUDA 15.363 / 15.601 ms 22.859 / 23.815 ms 47.541 / 47.831 ms
rrstitcher / OpenGL 14.923 / 15.282 ms 23.656 / 24.526 ms 51.661 / 52.775 ms
rrglstitcher / OpenGL 9.310 / 9.583 ms 18.834 / 20.349 ms 39.585 / 41.976 ms

These values represent the latency added between the Stitcher input boundary and stitcher.src. They do not include camera capture, decoding, display, or glass-to-glass latency.

For rrglstitcher values also exclude glupload, since the input marker is placed after the upload stage.


Note
The latency benchmark uses unpaced synthetic sources. The 30/1 framerate is part of the negotiated caps and does not rate-limit the sources to 30 FPS.


For GL-memory paths, observing a buffer at stitcher.src does not necessarily mean that all queued GPU commands have completed. The latency values should therefore be compared within the same benchmark method and workload.

Reproducing the Results

The following calibrations and pipelines can be used to reproduce the benchmark workloads.

Reference Calibrations

To keep this page compact, only the three-input reference calibrations are included below. The 2-, 4-, and 6-input calibrations are used for the corresponding scaling results but are not reproduced inline.

1280x720 — cal_720p.json

{
    "homographies": [
        {
            "images": {
                "target": 1,
                "reference": 0
            },
            "matrix": {
                "h00": 1.6871646644830918,
                "h01": 0.05384674276575977,
                "h02": -412.4132501160003,
                "h10": 0.21427657254032648,
                "h11": 1.3320852190430963,
                "h12": -94.09260636008055,
                "h20": 0.0005781267862696429,
                "h21": 1.2328156473235378e-06,
                "h22": 1.0
            }
        },
        {
            "images": {
                "target": 2,
                "reference": 0
            },
            "matrix": {
                "h00": 0.5161457595374556,
                "h01": -0.10346458895613274,
                "h02": 319.9377286523319,
                "h10": -0.10919438044207914,
                "h11": 0.7858616051182961,
                "h12": -8.496359451449297,
                "h20": -0.0003690709230813509,
                "h21": -5.038338091456794e-06,
                "h22": 1.0
            }
        }
    ]
}

1920x1080 — cal.json

{
    "homographies": [
        {
            "images": {
                "target": 1,
                "reference": 0
            },
            "matrix": {
                "h00": 1.6871646644830918,
                "h01": 0.05384674276575977,
                "h02": -618.6198751740005,
                "h10": 0.21427657254032648,
                "h11": 1.3320852190430963,
                "h12": -141.13890954012084,
                "h20": 0.0003854178575130953,
                "h21": 8.218770982156919e-07,
                "h22": 1.0
            }
        },
        {
            "images": {
                "target": 2,
                "reference": 0
            },
            "matrix": {
                "h00": 0.5161457595374556,
                "h01": -0.10346458895613274,
                "h02": 479.9065929784979,
                "h10": -0.10919438044207914,
                "h11": 0.7858616051182961,
                "h12": -12.744539177173948,
                "h20": -0.0002460472820542339,
                "h21": -3.358892060971196e-06,
                "h22": 1.0
            }
        }
    ]
}

3840x2160 — cal_4k.json

{
    "homographies": [
        {
            "images": {
                "target": 1,
                "reference": 0
            },
            "matrix": {
                "h00": 1.6871646644830918,
                "h01": 0.05384674276575977,
                "h02": -1237.239750348001,
                "h10": 0.21427657254032648,
                "h11": 1.3320852190430963,
                "h12": -282.2778190802417,
                "h20": 0.00019270892875654765,
                "h21": 4.1093854910784595e-07,
                "h22": 1.0
            }
        },
        {
            "images": {
                "target": 2,
                "reference": 0
            },
            "matrix": {
                "h00": 0.5161457595374556,
                "h01": -0.10346458895613274,
                "h02": 959.8131859569958,
                "h10": -0.10919438044207914,
                "h11": 0.7858616051182961,
                "h12": -25.489078354347896,
                "h20": -0.00012302364102711695,
                "h21": -1.679446030485598e-06,
                "h22": 1.0
            }
        }
    ]
}

Synthetic Benchmark Pipelines

Set the calibration and number of frames before running the test:

export CALIBRATION_FILE=cal.json
export FRAMES=300

Use FRAMES=300 for DragonWing and i.MX95, and FRAMES=600 for Jetson AGX Orin.

For the other input resolutions, use:

Resolution Calibration DragonWing i.MX95 Orin
1280x720 cal_720p.json 1500 frames 900 frames 3000 frames
1920x1080 cal.json 300 frames 300 frames 600 frames
3840x2160 cal_4k.json 300 frames 300 frames 600 frames

rrstitcher — three 1920x1080 synthetic inputs

gst-launch-1.0 -e \
    rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
    videotestsrc num-buffers="$FRAMES" pattern=black ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_0 \
    videotestsrc num-buffers="$FRAMES" pattern=black ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_1 \
    videotestsrc num-buffers="$FRAMES" pattern=black ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_2

The backend used by rrstitcher depends on the installed LibPanorama build. On Jetson, the CUDA and OpenGL results were collected using separate LibPanorama builds.

rrglstitcher — three 1920x1080 synthetic inputs

GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
gst-launch-1.0 -e \
    rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
    videotestsrc num-buffers="$FRAMES" pattern=black ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! \
        "video/x-raw(memory:GLMemory),format=RGBA" ! \
        queue ! stitcher.sink_0 \
    videotestsrc num-buffers="$FRAMES" pattern=black ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! \
        "video/x-raw(memory:GLMemory),format=RGBA" ! \
        queue ! stitcher.sink_1 \
    videotestsrc num-buffers="$FRAMES" pattern=black ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! \
        "video/x-raw(memory:GLMemory),format=RGBA" ! \
        queue ! stitcher.sink_2

Recorded-Video Benchmark Pipelines

The recorded-video tests use:

export VIDEO_0=/path/to/cam4_und.mp4
export VIDEO_1=/path/to/cam2_und.mp4
export VIDEO_2=/path/to/cam6_und.mp4
export CALIBRATION_FILE=/path/to/cal.json

DragonWing

DragonWing rrstitcher

gst-launch-1.0 -e \
    rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
    filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
    demux0.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
        "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_0 \
    filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
    demux1.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
        "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_1 \
    filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
    demux2.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
        "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_2

DragonWing rrglstitcher

GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
gst-launch-1.0 -e \
    rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
    filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
    demux0.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
        "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! "video/x-raw(memory:GLMemory),format=RGBA,texture-target=2D" ! \
        queue ! stitcher.sink_0 \
    filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
    demux1.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
        "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! "video/x-raw(memory:GLMemory),format=RGBA,texture-target=2D" ! \
        queue ! stitcher.sink_1 \
    filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
    demux2.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
        "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! "video/x-raw(memory:GLMemory),format=RGBA,texture-target=2D" ! \
        queue ! stitcher.sink_2

NXP i.MX95

i.MX95 rrstitcher

gst-launch-1.0 -e \
    rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! fakesink sync=false \
    filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
    demux0.video_0 ! queue ! h264parse ! v4l2h264dec ! \
        glupload ! glcolorconvert ! \
        "video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
        gldownload ! "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_0 \
    filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
    demux1.video_0 ! queue ! h264parse ! v4l2h264dec ! \
        glupload ! glcolorconvert ! \
        "video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
        gldownload ! "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_1 \
    filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
    demux2.video_0 ! queue ! h264parse ! v4l2h264dec ! \
        glupload ! glcolorconvert ! \
        "video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
        gldownload ! "video/x-raw,format=RGBA,width=1920,height=1080" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_2

i.MX95 rrglstitcher

GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
gst-launch-1.0 -e \
    rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! fakesink sync=false \
    filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
    demux0.video_0 ! queue ! h264parse ! v4l2h264dec ! \
        glupload ! glcolorconvert ! \
        "video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_0 \
    filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
    demux1.video_0 ! queue ! h264parse ! v4l2h264dec ! \
        glupload ! glcolorconvert ! \
        "video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_1 \
    filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
    demux2.video_0 ! queue ! h264parse ! v4l2h264dec ! \
        glupload ! glcolorconvert ! \
        "video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_2

NVIDIA Jetson AGX Orin

Orin rrstitcher

gst-launch-1.0 -e \
    rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
    filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
    demux0.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_0 \
    filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
    demux1.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_1 \
    filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
    demux2.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        stitcher.sink_2

Orin rrglstitcher

GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
GST_GL_WINDOW=surfaceless \
gst-launch-1.0 -e \
    rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
    stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
    filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
    demux0.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! "video/x-raw(memory:GLMemory),format=RGBA" ! \
        queue ! stitcher.sink_0 \
    filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
    demux1.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! "video/x-raw(memory:GLMemory),format=RGBA" ! \
        queue ! stitcher.sink_1 \
    filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
    demux2.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
        "video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
        queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
        glupload ! "video/x-raw(memory:GLMemory),format=RGBA" ! \
        queue ! stitcher.sink_2

Jetson OpenGL Notes

For rrglstitcher measurements on Jetson AGX Orin, the following GL configuration was used:

export GST_GL_PLATFORM=egl
export GST_GL_API=gles2
export GST_GL_WINDOW=surfaceless

OpenGL-backed rrstitcher creates its own LibPanorama OpenGL context. Depending on the graphical session, access to the active display may be required.

Other Resolutions and Input Counts

To reproduce a 720p or 4K synthetic result, change the calibration file and the caps in the pipeline to the corresponding resolution.

To reproduce the 2-, 4-, or 6-input synthetic results, use the calibration file for that input count and add or remove videotestsrc branches so that the inputs are connected contiguously from sink_0 to the last camera index.

Run one warm-up trial before collecting measurements, followed by at least three recorded runs. The system should otherwise remain idle.

Constraints

  • The four- and six-input benchmarks use synthetic extensions of the three-camera calibration. They are intended to measure scaling and do not represent physical four- or six-camera rigs.
  • Input count and panorama size increase together in these tests. Performance therefore depends on both the number of cameras and the output dimensions produced by the calibration.
  • Large 4K configurations are memory-limited on the tested i.MX95. The four-input GL-memory case was not stable, and the six-input system-memory case reached OOM.
  • Jetson AGX Orin throughput depends on the clock configuration. Results collected in MAXN and MAXN with jetson_clocks should be treated as separate configurations.




Cookies help us deliver our services. By using our services, you agree to our use of cookies.