Performance
🚧 Documentation is under development
The Video Stitching for Embedded Systems guide is currently under active development. Some sections may be incomplete or change without notice.
Questions? Contact RidgeRun or email to support@ridgerun.com.
Performance
This page contains the performance measurements collected for the Stitcher on the supported embedded platforms.
The main benchmarks use synthetic RGBA inputs so the measurements reflect stitching throughput without adding decoding, encoding, display, or file I/O to the measured path.
Benchmark Methodology
Test Workload
The synthetic benchmarks use videotestsrc to feed RGBA frames directly into the Stitcher. The output is discarded with fakesink sync=false.
The tested input resolutions are:
- 1280x720
- 1920x1080
- 3840x2160
The tested input counts are:
- 2 cameras
- 3 cameras
- 4 cameras
- 6 cameras
All inputs use RGBA and 30/1 framerate caps.
The sources are non-live and the output sink is not synchronized, so the benchmark is not limited to 30 FPS. Values above 30 FPS mean the pipeline can process frames faster than real time for that configuration.
The system-memory path uses rrstitcher:
videotestsrc ! \
video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1 ! \
queue ! stitcher.sink_N
The GL-memory path uses rrglstitcher:
videotestsrc ! \
video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1 ! \
queue ! glupload ! \
video/x-raw(memory:GLMemory),format=RGBA ! \
queue ! stitcher.sink_N
The glupload stage is required to provide GstGLMemory to rrglstitcher.
Calibration and Camera Layout
The same calibration geometry was used on all platforms for a given resolution and camera count.
The three-camera 1920x1080 calibration is the reference case. The 720p and 4K calibrations use the same geometry with the homographies scaled to the corresponding input resolution.
The two-camera case uses cameras 0 and 1 from the reference calibration.
The four- and six-camera benchmarks use synthetic extensions of the three-camera calibration. Additional camera positions are created by translating the last homography. These cases measure scaling with additional inputs and larger panoramas; they are not physical four- or six-camera calibrations.
The resulting panorama dimensions are:
| Inputs | 1280x720 | 1920x1080 | 3840x2160 |
|---|---|---|---|
| 2 | 1693x959 | 2539x1438 | 5078x2875 |
| 3 | 2270x1145 | 3406x1718 | 6812x3436 |
| 4 | 2849x1145 | 4274x1718 | 8547x3436 |
| 6 | 4005x1145 | 6008x1718 | 12015x3436 |
Overlap is defined by the calibration. A fixed overlap percentage was not used as part of this benchmark.
Measurement Method
Each configuration was warmed up once before collecting results. Recorded configurations were then run at least three times sequentially.
A run was only accepted if it processed all requested frames and reached EOS without an error.
FPS was obtained from the final mean_fps reported by the GStreamer perf element where available.
CPU and RAM values are whole-system measurements, not Stitcher process usage.
For Jetson AGX Orin, CPU, RAM, and GPU usage were collected with tegrastats at 500 ms intervals.
For DragonWing and i.MX95, CPU and RAM were sampled from /proc/stat and /proc/meminfo every 500 ms.
DragonWing GPU utilization was obtained from the KGSL device-wide busy counter.
A reliable GPU utilization counter was not available on i.MX95, so GPU usage is marked as unavailable.
The test system was kept otherwise idle while collecting each result.
Recorded-Video Measurement
A smaller recorded-video benchmark was performed to compare the synthetic three-input 1920x1080 results against pipelines using actual camera recordings.
The same three recordings and three-camera calibration are used throughout the comparison. The files are read with filesrc and the output is discarded with fakesink sync=false, so the pipeline runs as fast as possible rather than at the normal playback rate of the recordings.
The test does not include output encoding, display, muxing, or file output.
The input processing path depends on the platform:
| Platform | Decode and conversion path |
|---|---|
| NVIDIA Jetson AGX Orin | Hardware H.264 decoding with nvv4l2decoder and conversion to RGBA with nvvidconv.
|
| Qualcomm DragonWing | Software H.264 decoding with avdec_h264 and videoconvert.
|
| NXP i.MX95 | Hardware H.264 decoding with v4l2h264dec, followed by glupload and glcolorconvert. rrstitcher additionally uses gldownload before its system-memory input.
|
The recorded-video tests are intended as a practical comparison with the synthetic workload. They are not a separate camera-count or resolution performance matrix.
Latency Measurement Method
Latency was measured separately from the throughput benchmark using 1920x1080 RGBA synthetic inputs with 30/1 caps. The latency benchmark covers 2, 3, and 6 inputs.
For rrstitcher, an identity marker is placed immediately before each Stitcher sink. For rrglstitcher, the marker is placed after glupload, so the GL upload is not included in the reported latency.
A probe at each input marker and a probe at stitcher.src record a monotonic timestamp when each buffer passes. Input and output buffers are matched using their PTS.
For each output frame, latency is calculated as the largest input-to-output delay across all camera inputs:
Latency = max(output probe time - input marker probe time)
Tested Platform Configurations
The benchmark used the following platform configurations:
| Platform | Stitcher configuration | Test conditions |
|---|---|---|
| NVIDIA Jetson AGX Orin | rrstitcher with separate CUDA and OpenGL LibPanorama builds; rrglstitcher with OpenGL
|
MAXN power mode. Tested with dynamic clocks and with jetson_clocks.
|
| Qualcomm DragonWing | rrstitcher with OpenGL; rrglstitcher with OpenGL
|
CPU/RAM sampled from /proc. GPU sampled through KGSL.
|
| NXP i.MX95 | rrstitcher with OpenGL; rrglstitcher with OpenGL
|
Active Weston session. CPU/RAM sampled from /proc. GPU usage unavailable.
|
For the Orin jetson_clocks runs, the CPU cores operated at approximately 2201.6 MHz, the GPU at 1300.5 MHz, and EMC at 3199 MHz.
The Orin latency measurements were collected in MAXN mode with jetson_clocks enabled.
Results by Platform and Configuration
FPS values are rounded to the nearest whole number in the main synthetic throughput tables below.
30 FPS Capability Summary
The table below shows the highest tested camera count that reached at least 30 FPS for each input resolution.
| Platform | Element / Backend | 1280x720 | 1920x1080 | 3840x2160 |
|---|---|---|---|---|
| DragonWing | rrstitcher / OpenGL
|
6* | 4 | — |
rrglstitcher / OpenGL
|
6* | 4 | 2 | |
| i.MX95 | rrstitcher / OpenGL
|
2 | — | — |
rrglstitcher / OpenGL
|
4 | 3 | — | |
| Orin AGX MAXN | rrstitcher / CUDA
|
6* | 6* | 2 |
rrstitcher / OpenGL
|
6* | 6* | 2 | |
rrglstitcher / OpenGL
|
6* | 6* | 3 | |
Orin AGX MAXN + jetson_clocks
|
rrstitcher / CUDA
|
6* | 6* | 3 |
rrstitcher / OpenGL
|
6* | 6* | 3 | |
rrglstitcher / OpenGL
|
6* | 6* | 3 |
* Six cameras was the highest input count tested. These entries should not be interpreted as a six-camera product limit.
A dash means that the two-camera case did not reach 30 FPS at that resolution.
This table can also be used to estimate the highest tested resolution for a required camera count. For example, the DragonWing rrglstitcher reached 30 FPS with two 4K inputs, while four inputs reached 30 FPS up to 1080p.
Qualcomm DragonWing
| Resolution | Inputs | rrstitcher
|
rrglstitcher
|
|---|---|---|---|
| 1280x720 | 2 | 131 FPS | 109 FPS |
| 3 | 104 FPS | 80 FPS | |
| 4 | 90 FPS | 74 FPS | |
| 6 | 61 FPS | 50 FPS | |
| 1920x1080 | 2 | 78 FPS | 96 FPS |
| 3 | 49 FPS | 57 FPS | |
| 4 | 37 FPS | 45 FPS | |
| 6 | 22 FPS | 27 FPS | |
| 3840x2160 | 2 | 27 FPS | 41 FPS |
| 3 | 15 FPS | 24 FPS | |
| 4 | 13 FPS | 20 FPS | |
| 6 | 7 FPS | 12 FPS |
NXP i.MX95
| Resolution | Inputs | rrstitcher
|
rrglstitcher
|
|---|---|---|---|
| 1280x720 | 2 | 52 FPS | 85 FPS |
| 3 | 29 FPS | 42 FPS | |
| 4 | 20 FPS | 36 FPS | |
| 6 | 14 FPS | 21 FPS | |
| 1920x1080 | 2 | 25 FPS | 69 FPS |
| 3 | 15 FPS | 34 FPS | |
| 4 | 9 FPS | 22 FPS | |
| 6 | 7 FPS | 11 FPS | |
| 3840x2160 | 2 | 7 FPS | 20 FPS |
| 3 | 4 FPS | 10 FPS | |
| 4 | 3 FPS | Not measurable | |
| 6 | Unsupported (OOM) | Not tested |
NVIDIA Jetson AGX Orin
MAXN - Dynamic Clocks
| Resolution | Inputs | rrstitcher CUDA
|
rrstitcher OpenGL
|
rrglstitcher
|
|---|---|---|---|---|
| 1280x720 | 2 | 211 FPS | 278 FPS | 294 FPS |
| 3 | 150 FPS | 157 FPS | 149 FPS | |
| 4 | 115 FPS | 112 FPS | 121 FPS | |
| 6 | 77 FPS | 71 FPS | 72 FPS | |
| 1920x1080 | 2 | 109 FPS | 133 FPS | 171 FPS |
| 3 | 68 FPS | 85 FPS | 99 FPS | |
| 4 | 55 FPS | 57 FPS | 80 FPS | |
| 6 | 36 FPS | 37 FPS | 58 FPS | |
| 3840x2160 | 2 | 32 FPS | 41 FPS | 51 FPS |
| 3 | 21 FPS | 28 FPS | 48 FPS | |
| 4 | 18 FPS | 21 FPS | 21 FPS | |
| 6 | 13 FPS | 14 FPS | 14 FPS |
MAXN with jetson_clocks
| Resolution | Inputs | rrstitcher CUDA
|
rrstitcher OpenGL
|
rrglstitcher
|
|---|---|---|---|---|
| 1280x720 | 2 | 437 FPS | 401 FPS | 618 FPS |
| 3 | 280 FPS | 262 FPS | 367 FPS | |
| 4 | 213 FPS | 199 FPS | 311 FPS | |
| 6 | 139 FPS | 126 FPS | 180 FPS | |
| 1920x1080 | 2 | 200 FPS | 197 FPS | 248 FPS |
| 3 | 128 FPS | 125 FPS | 205 FPS | |
| 4 | 96 FPS | 93 FPS | 120 FPS | |
| 6 | 62 FPS | 56 FPS | 86 FPS | |
| 3840x2160 | 2 | 53 FPS | 54 FPS | 64 FPS |
| 3 | 33 FPS | 33 FPS | 57 FPS | |
| 4 | 25 FPS | 24 FPS | 28 FPS | |
| 6 | 16 FPS | 15 FPS | 20 FPS |
Synthetic vs Recorded Video
The table below compares the synthetic three-input 1920x1080 workload with the recorded-video pipelines.
The synthetic result isolates the Stitcher, while the recorded-video result also includes the platform's decode and input-conversion path.
| Platform / Configuration | Element / Backend | Decode | Synthetic | Recorded video |
|---|---|---|---|---|
| DragonWing | rrstitcher / OpenGL
|
Software | 49 FPS | 52 FPS |
rrglstitcher / OpenGL
|
Software | 57 FPS | 63 FPS | |
| i.MX95 | rrstitcher / OpenGL
|
Hardware | 15 FPS | 16 FPS |
rrglstitcher / OpenGL
|
Hardware | 34 FPS | 26 FPS | |
| Orin AGX MAXN | rrstitcher / CUDA
|
Hardware | 68 FPS | 88 FPS |
rrstitcher / OpenGL
|
Hardware | 85 FPS | 93 FPS | |
rrglstitcher / OpenGL
|
Hardware | 99 FPS | 117 FPS | |
Orin AGX MAXN + jetson_clocks
|
rrstitcher / CUDA
|
Hardware | 128 FPS | 123 FPS |
rrstitcher / OpenGL
|
Hardware | 125 FPS | 122 FPS | |
rrglstitcher / OpenGL
|
Hardware | 205 FPS | 124 FPS |
Resource Usage - Three 1920x1080 Synthetic Inputs
This table shows resource usage for the three-input 1080p synthetic benchmark.
These measurements were collected while the pipelines were running at maximum throughput. They do not represent a pipeline rate-limited to 30 FPS.
| Platform | Element / Backend | FPS | CPU Avg. | RAM Avg. / Peak | GPU Avg. / Peak |
|---|---|---|---|---|---|
| DragonWing | rrstitcher / OpenGL
|
49 | 14.50% | 5.06% / 5.44% | 43.02% / 48.68% |
| DragonWing | rrglstitcher / OpenGL
|
57 | 13.55% | 5.52% / 6.36% | 47.91% / 56.72% |
| i.MX95 | rrstitcher / OpenGL
|
15 | 16.01% | 18.49% / 20.07% | Unavailable |
| i.MX95 | rrglstitcher / OpenGL
|
34 | 18.51% | 17.21% / 21.60% | Unavailable |
| Orin AGX MAXN | rrstitcher / CUDA
|
68 | 8.64% | 6.40% / 6.52% | 55.14% / 88% |
| Orin AGX MAXN | rrstitcher / OpenGL
|
85 | 9.54% | 9.41% / 9.47% | 54.00% / 84% |
| Orin AGX MAXN | rrglstitcher / OpenGL
|
99 | 15.13% | 6.17% / 6.58% | 47.60% / 89% |
Orin AGX MAXN + jetson_clocks
|
rrstitcher / CUDA
|
128 | 9.88% | 8.55% / 9.28% | 28.14% / 50% |
Orin AGX MAXN + jetson_clocks
|
rrstitcher / OpenGL
|
125 | 11.34% | 9.41% / 9.51% | 34.53% / 59% |
Orin AGX MAXN + jetson_clocks
|
rrglstitcher / OpenGL
|
205 | 15.73% | 8.32% / 9.07% | 43.17% / 78% |
CPU and RAM values include the rest of the system. GPU values are device-wide and may include work outside the Stitcher.
Latency
The following table shows the measured Stitcher input-to-output latency for 1920x1080 RGBA synthetic inputs.
Values are shown as median / p95 in milliseconds.
| Platform | Element / Backend | 2 Inputs | 3 Inputs | 6 Inputs |
|---|---|---|---|---|
| DragonWing | rrstitcher / OpenGL
|
33.340 / 45.404 ms | 54.601 / 72.864 ms | 122.425 / 165.895 ms |
rrglstitcher / OpenGL
|
23.353 / 28.123 ms | 41.104 / 50.244 ms | 101.083 / 115.283 ms | |
| i.MX95 | rrstitcher / OpenGL
|
125.350 / 138.248 ms | 221.081 / 229.312 ms | 450.882 / 454.592 ms |
rrglstitcher / OpenGL
|
41.425 / 43.229 ms | 85.984 / 89.472 ms | 272.884 / 278.165 ms | |
Orin AGX MAXN + jetson_clocks
|
rrstitcher / CUDA
|
15.363 / 15.601 ms | 22.859 / 23.815 ms | 47.541 / 47.831 ms |
rrstitcher / OpenGL
|
14.923 / 15.282 ms | 23.656 / 24.526 ms | 51.661 / 52.775 ms | |
rrglstitcher / OpenGL
|
9.310 / 9.583 ms | 18.834 / 20.349 ms | 39.585 / 41.976 ms |
These values represent the latency added between the Stitcher input boundary and stitcher.src. They do not include camera capture, decoding, display, or glass-to-glass latency.
For rrglstitcher values also exclude glupload, since the input marker is placed after the upload stage.
For GL-memory paths, observing a buffer at stitcher.src does not necessarily mean that all queued GPU commands have completed. The latency values should therefore be compared within the same benchmark method and workload.
Reproducing the Results
The following calibrations and pipelines can be used to reproduce the benchmark workloads.
Reference Calibrations
To keep this page compact, only the three-input reference calibrations are included below. The 2-, 4-, and 6-input calibrations are used for the corresponding scaling results but are not reproduced inline.
1280x720 — cal_720p.json
{
"homographies": [
{
"images": {
"target": 1,
"reference": 0
},
"matrix": {
"h00": 1.6871646644830918,
"h01": 0.05384674276575977,
"h02": -412.4132501160003,
"h10": 0.21427657254032648,
"h11": 1.3320852190430963,
"h12": -94.09260636008055,
"h20": 0.0005781267862696429,
"h21": 1.2328156473235378e-06,
"h22": 1.0
}
},
{
"images": {
"target": 2,
"reference": 0
},
"matrix": {
"h00": 0.5161457595374556,
"h01": -0.10346458895613274,
"h02": 319.9377286523319,
"h10": -0.10919438044207914,
"h11": 0.7858616051182961,
"h12": -8.496359451449297,
"h20": -0.0003690709230813509,
"h21": -5.038338091456794e-06,
"h22": 1.0
}
}
]
}
1920x1080 — cal.json
{
"homographies": [
{
"images": {
"target": 1,
"reference": 0
},
"matrix": {
"h00": 1.6871646644830918,
"h01": 0.05384674276575977,
"h02": -618.6198751740005,
"h10": 0.21427657254032648,
"h11": 1.3320852190430963,
"h12": -141.13890954012084,
"h20": 0.0003854178575130953,
"h21": 8.218770982156919e-07,
"h22": 1.0
}
},
{
"images": {
"target": 2,
"reference": 0
},
"matrix": {
"h00": 0.5161457595374556,
"h01": -0.10346458895613274,
"h02": 479.9065929784979,
"h10": -0.10919438044207914,
"h11": 0.7858616051182961,
"h12": -12.744539177173948,
"h20": -0.0002460472820542339,
"h21": -3.358892060971196e-06,
"h22": 1.0
}
}
]
}
3840x2160 — cal_4k.json
{
"homographies": [
{
"images": {
"target": 1,
"reference": 0
},
"matrix": {
"h00": 1.6871646644830918,
"h01": 0.05384674276575977,
"h02": -1237.239750348001,
"h10": 0.21427657254032648,
"h11": 1.3320852190430963,
"h12": -282.2778190802417,
"h20": 0.00019270892875654765,
"h21": 4.1093854910784595e-07,
"h22": 1.0
}
},
{
"images": {
"target": 2,
"reference": 0
},
"matrix": {
"h00": 0.5161457595374556,
"h01": -0.10346458895613274,
"h02": 959.8131859569958,
"h10": -0.10919438044207914,
"h11": 0.7858616051182961,
"h12": -25.489078354347896,
"h20": -0.00012302364102711695,
"h21": -1.679446030485598e-06,
"h22": 1.0
}
}
]
}
Synthetic Benchmark Pipelines
Set the calibration and number of frames before running the test:
export CALIBRATION_FILE=cal.json export FRAMES=300
Use FRAMES=300 for DragonWing and i.MX95, and FRAMES=600 for Jetson AGX Orin.
For the other input resolutions, use:
| Resolution | Calibration | DragonWing | i.MX95 | Orin |
|---|---|---|---|---|
| 1280x720 | cal_720p.json
|
1500 frames | 900 frames | 3000 frames |
| 1920x1080 | cal.json
|
300 frames | 300 frames | 600 frames |
| 3840x2160 | cal_4k.json
|
300 frames | 300 frames | 600 frames |
rrstitcher — three 1920x1080 synthetic inputs
gst-launch-1.0 -e \
rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
videotestsrc num-buffers="$FRAMES" pattern=black ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_0 \
videotestsrc num-buffers="$FRAMES" pattern=black ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_1 \
videotestsrc num-buffers="$FRAMES" pattern=black ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_2
The backend used by rrstitcher depends on the installed LibPanorama build. On Jetson, the CUDA and OpenGL results were collected using separate LibPanorama builds.
rrglstitcher — three 1920x1080 synthetic inputs
GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
gst-launch-1.0 -e \
rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
videotestsrc num-buffers="$FRAMES" pattern=black ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! \
"video/x-raw(memory:GLMemory),format=RGBA" ! \
queue ! stitcher.sink_0 \
videotestsrc num-buffers="$FRAMES" pattern=black ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! \
"video/x-raw(memory:GLMemory),format=RGBA" ! \
queue ! stitcher.sink_1 \
videotestsrc num-buffers="$FRAMES" pattern=black ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! \
"video/x-raw(memory:GLMemory),format=RGBA" ! \
queue ! stitcher.sink_2
Recorded-Video Benchmark Pipelines
The recorded-video tests use:
export VIDEO_0=/path/to/cam4_und.mp4 export VIDEO_1=/path/to/cam2_und.mp4 export VIDEO_2=/path/to/cam6_und.mp4 export CALIBRATION_FILE=/path/to/cal.json
DragonWing
DragonWing rrstitcher
gst-launch-1.0 -e \
rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
demux0.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
"video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_0 \
filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
demux1.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
"video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_1 \
filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
demux2.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
"video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_2
DragonWing rrglstitcher
GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
gst-launch-1.0 -e \
rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
demux0.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
"video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! "video/x-raw(memory:GLMemory),format=RGBA,texture-target=2D" ! \
queue ! stitcher.sink_0 \
filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
demux1.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
"video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! "video/x-raw(memory:GLMemory),format=RGBA,texture-target=2D" ! \
queue ! stitcher.sink_1 \
filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
demux2.video_0 ! queue ! h264parse ! avdec_h264 ! videoconvert ! \
"video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! "video/x-raw(memory:GLMemory),format=RGBA,texture-target=2D" ! \
queue ! stitcher.sink_2
NXP i.MX95
i.MX95 rrstitcher
gst-launch-1.0 -e \
rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! fakesink sync=false \
filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
demux0.video_0 ! queue ! h264parse ! v4l2h264dec ! \
glupload ! glcolorconvert ! \
"video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
gldownload ! "video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_0 \
filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
demux1.video_0 ! queue ! h264parse ! v4l2h264dec ! \
glupload ! glcolorconvert ! \
"video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
gldownload ! "video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_1 \
filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
demux2.video_0 ! queue ! h264parse ! v4l2h264dec ! \
glupload ! glcolorconvert ! \
"video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
gldownload ! "video/x-raw,format=RGBA,width=1920,height=1080" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_2
i.MX95 rrglstitcher
GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
gst-launch-1.0 -e \
rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! fakesink sync=false \
filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
demux0.video_0 ! queue ! h264parse ! v4l2h264dec ! \
glupload ! glcolorconvert ! \
"video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_0 \
filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
demux1.video_0 ! queue ! h264parse ! v4l2h264dec ! \
glupload ! glcolorconvert ! \
"video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_1 \
filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
demux2.video_0 ! queue ! h264parse ! v4l2h264dec ! \
glupload ! glcolorconvert ! \
"video/x-raw(memory:GLMemory),format=RGBA,width=1920,height=1080,texture-target=2D" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_2
NVIDIA Jetson AGX Orin
Orin rrstitcher
gst-launch-1.0 -e \
rrstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
demux0.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_0 \
filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
demux1.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_1 \
filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
demux2.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
stitcher.sink_2
Orin rrglstitcher
GST_GL_PLATFORM=egl \
GST_GL_API=gles2 \
GST_GL_WINDOW=surfaceless \
gst-launch-1.0 -e \
rrglstitcher name=stitcher calibration-file="$CALIBRATION_FILE" \
stitcher. ! queue ! perf name=measure print-cpu-load=true ! fakesink sync=false \
filesrc location="$VIDEO_0" ! qtdemux name=demux0 \
demux0.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! "video/x-raw(memory:GLMemory),format=RGBA" ! \
queue ! stitcher.sink_0 \
filesrc location="$VIDEO_1" ! qtdemux name=demux1 \
demux1.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! "video/x-raw(memory:GLMemory),format=RGBA" ! \
queue ! stitcher.sink_1 \
filesrc location="$VIDEO_2" ! qtdemux name=demux2 \
demux2.video_0 ! queue ! h264parse ! nvv4l2decoder ! nvvidconv ! \
"video/x-raw,format=RGBA,width=1920,height=1080,framerate=30/1" ! \
queue max-size-buffers=2 max-size-bytes=0 max-size-time=0 ! \
glupload ! "video/x-raw(memory:GLMemory),format=RGBA" ! \
queue ! stitcher.sink_2
Jetson OpenGL Notes
For rrglstitcher measurements on Jetson AGX Orin, the following GL configuration was used:
export GST_GL_PLATFORM=egl export GST_GL_API=gles2 export GST_GL_WINDOW=surfaceless
OpenGL-backed rrstitcher creates its own LibPanorama OpenGL context. Depending on the graphical session, access to the active display may be required.
Other Resolutions and Input Counts
To reproduce a 720p or 4K synthetic result, change the calibration file and the caps in the pipeline to the corresponding resolution.
To reproduce the 2-, 4-, or 6-input synthetic results, use the calibration file for that input count and add or remove videotestsrc branches so that the inputs are connected contiguously from sink_0 to the last camera index.
Run one warm-up trial before collecting measurements, followed by at least three recorded runs. The system should otherwise remain idle.
Constraints
- The four- and six-input benchmarks use synthetic extensions of the three-camera calibration. They are intended to measure scaling and do not represent physical four- or six-camera rigs.
- Input count and panorama size increase together in these tests. Performance therefore depends on both the number of cameras and the output dimensions produced by the calibration.
- Large 4K configurations are memory-limited on the tested i.MX95. The four-input GL-memory case was not stable, and the six-input system-memory case reached OOM.
- Jetson AGX Orin throughput depends on the clock configuration. Results collected in MAXN and MAXN with
jetson_clocksshould be treated as separate configurations.