TL;DR. A camera is a sampling operator that maps each finite pixel to a bundle of rays. ConeGaussian builds an anisotropic pixel footprint directly from the neighbouring rays of any calibrated central camera's inverse projection, filters the exact ray–Gaussian response with it in closed form, and derives a per-Gaussian training-frequency floor from the same geometry. One implementation gives anti-aliased zoom-out and artifact-free zoom-in for pinhole and strongly distorted fisheye cameras, and ports unchanged across two Gaussian ray-rendering backbones (3DGEER and 3DGUT).
In rendering, a camera is a sampling operator that maps each finite pixel to a bundle of rays. Different camera models change the geometry of this bundle, thus making a unified and faithful rendering formulation challenging. Consequently, Gaussian ray tracing supports generic cameras (with optical center) through their inverse ray mappings, yet typically reduces every pixel to a single center ray. This ignores the camera-dependent pixel footprint, causing aliasing under minification, while unconstrained Gaussians expose unsupported frequencies under magnification. We present ConeGaussian, a camera-model-agnostic anti-aliasing framework for Gaussian ray-based rendering. Instead of defining the pixel filter on a camera-specific image plane, ConeGaussian constructs an anisotropic footprint directly from neighboring rays produced by the camera's native inverse mapping. We derive a closed-form response under a locally linear, depth-local, moment-matched approximation of the finite pixel footprint, while the same geometry defines a per-Gaussian training-frequency floor. Notably, by construction, our filtering principle can be used unmodified across calibrated central camera models and multiple Gaussian ray-rendering backbones. Additionally, unlike in Mip-Splatting, our scene-space frequency floor and filtering enable trivial composition at render time, allowing us to remove excess blurring. On pinhole and strongly distorted fisheye captures, ConeGaussian consistently improves two distinct ray-based backbones, by up to 4.3 dB at 1/8 resolution, and reduces fisheye LPIPS by 30% where perspective screen-plane footprint formulations are not directly applicable.
Move the pixel across the sensor, switch the camera model or reshape the Gaussian. The 3D geometry and both rows of filters update together.















Choose Zoom-out or Zoom-in, then a scene, to compare 3DGEER → ConeGaussian → ground truth. The 6 zoom-out and 5 zoom-in comparisons are kept separate, with their original detail insets and scale labels. Scene/view names are descriptive because the source figures do not list scene IDs.
Zoom-out · 1/8 sampling rate. Select a scene below.



Supplementary zoomout figure, row 6. Full source figure ↗



Supplementary zoomout figure, row 1. Full source figure ↗



Supplementary zoomout figure, row 2. Full source figure ↗



Supplementary zoomout figure, row 3. Full source figure ↗



Supplementary zoomout figure, row 4. Full source figure ↗



Supplementary zoomout figure, row 5. Full source figure ↗
Zoom-in · ×8. Select a scene below.



Supplementary zoomin figure, row 5. Full source figure ↗



Supplementary zoomin figure, row 1. Full source figure ↗



Supplementary zoomin figure, row 2. Full source figure ↗



Supplementary zoomin figure, row 3. Full source figure ↗



Supplementary zoomin figure, row 4. Full source figure ↗
Every cell reads PSNR↑ / SSIM↑ / LPIPS↓. Real-scene benchmarks: ScanNet++ and Zip-NeRF (fisheye, metrics inside a 3-pixel-eroded valid-domain mask) and Mip-NeRF 360 (pinhole); every eighth image is held out. Zoom-in / zoom-out change only the image sampling resolution with consistently scaled calibration, never the optical zoom or the pose. Green = best, orange = second best per column and metric among the generalizable ray-based methods.
| Dataset | Method | 1 | ½ | ¼ | ⅛ | Avg. |
|---|---|---|---|---|---|---|
| ScanNet++fisheye | 3DGUT | 29.33/.915/.248 | 30.38/.934/.193 | 31.47/.954/.112 | 30.24/.953/.078 | 30.36/.939/.158 |
| ConeGaussian w. 3DGUT | 29.67/.921/.239 | 30.61/.938/.185 | 31.66/.956/.108 | 31.65/.963/.059 | 30.90/.944/.148 | |
| 3DGEER | 27.59/.911/.257 | 28.53/.929/.202 | 28.53/.946/.123 | 28.02/.946/.084 | 28.17/.933/.167 | |
| ConeGaussian w. 3DGEER | 27.65/.917/.243 | 28.29/.931/.189 | 29.27/.950/.111 | 29.93/.962/.052 | 28.78/.940/.149 | |
| Zip-NeRFfisheye | 3DGUT | 23.40/.781/.413 | 24.18/.835/.299 | 24.73/.867/.192 | 24.58/.865/.147 | 24.22/.837/.263 |
| ConeGaussian w. 3DGUT | 23.48/.791/.403 | 24.22/.840/.292 | 24.82/.871/.185 | 25.14/.880/.128 | 24.42/.845/.252 | |
| 3DGEER | 24.32/.806/.388 | 25.02/.854/.268 | 25.64/.893/.162 | 25.45/.895/.121 | 25.11/.862/.235 | |
| ConeGaussian w. 3DGEER | 24.42/.815/.377 | 25.16/.860/.261 | 25.89/.898/.153 | 26.37/.915/.094 | 25.46/.872/.221 | |
| Mip-NeRF 360Pinhole | 3DGS | 26.55/.779/.274 | 28.00/.854/.162 | 28.51/.891/.102 | 27.45/.888/.087 | 27.63/.853/.156 |
| Mip-Splatting† | 27.20/.802/.244 | 28.74/.870/.146 | 29.90/.915/.090 | 30.66/.944/.056 | 29.12/.883/.134 | |
| Analytic-Splatting† | 27.50/.808/.231 | 28.99/.874/.132 | 30.35/.919/.077 | 31.21/.945/.051 | 29.51/.887/.123 | |
| 3DGUT | 26.68/.768/.341 | 26.80/.788/.241 | 26.97/.823/.159 | 24.02/.716/.238 | 26.12/.774/.245 | |
| ConeGaussian w. 3DGUT | 26.98/.782/.332 | 26.94/.798/.234 | 27.21/.831/.152 | 25.00/.762/.210 | 26.53/.793/.232 | |
| 3DGEER | 26.01/.739/.326 | 27.55/.820/.212 | 28.56/.880/.128 | 26.74/.860/.128 | 27.21/.825/.199 | |
| ConeGaussian w. 3DGEER | 27.05/.792/.275 | 28.44/.841/.176 | 30.03/.905/.095 | 31.08/.945/.052 | 29.15/.871/.149 |
Multi-scale training and testing at {1, ½, ¼, ⅛}. PSNR↑ / SSIM↑ / LPIPS↓. Gray rows are author-reported, pinhole-only screen-space methods (3DGS rasterizer) that do not support generic cameras; green / orange mark the best / second-best generalizable ray-based method per column and metric.
| Dataset | Method | 1 | ½ | ¼ | ⅛ | Avg. |
|---|---|---|---|---|---|---|
| ScanNet++fisheye | 3DGEER (baseline) | 27.89/.921/.243 | 28.18/.930/.198 | 27.10/.930/.155 | 26.58/.918/.131 | 27.44/.925/.182 |
| VKRayGS z/f | 26.20/.919/.244 | 25.89/.922/.205 | 21.57/.885/.193 | 20.54/.833/.185 | 23.55/.890/.207 | |
| ConeGaussian (isotropic) | 27.96/.921/.242 | 28.45/.932/.194 | 27.34/.937/.137 | 27.37/.938/.089 | 27.78/.932/.165 | |
| ConeGaussian (anisotropic) | 27.97/.922/.239 | 28.45/.932/.193 | 28.25/.942/.128 | 27.93/.942/.079 | 28.15/.934/.160 | |
| Zip-NeRFfisheye | 3DGEER (baseline) | 24.13/.787/.409 | 24.92/.845/.278 | 25.87/.899/.150 | 25.82/.906/.104 | 25.19/.859/.235 |
| VKRayGS z/f | 23.37/.786/.406 | 24.61/.846/.276 | 26.00/.902/.148 | 25.92/.913/.092 | 24.97/.862/.230 | |
| ConeGaussian (isotropic) | 24.16/.794/.401 | 25.01/.850/.271 | 25.97/.901/.148 | 26.41/.918/.089 | 25.39/.866/.227 | |
| ConeGaussian (anisotropic) | 24.25/.795/.398 | 25.02/.851/.270 | 25.96/.901/.148 | 26.43/.919/.088 | 25.41/.867/.226 | |
| Mip-NeRF 360pinhole | 3DGEER (baseline) | 27.03/.797/.300 | 27.68/.832/.214 | 27.19/.836/.171 | 25.29/.784/.183 | 26.80/.812/.217 |
| VKRayGS z/f | 27.14/.798/.298 | 27.25/.832/.219 | 25.68/.839/.175 | 22.96/.786/.183 | 25.76/.814/.219 | |
| ConeGaussian (isotropic) | 27.17/.799/.296 | 28.10/.843/.207 | 28.80/.884/.135 | 28.53/.903/.096 | 28.15/.857/.183 | |
| ConeGaussian (anisotropic) | 27.11/.803/.296 | 28.06/.829/.206 | 28.87/.876/.131 | 28.67/.909/.087 | 28.18/.854/.180 |
Zoom-out (STMT): trained at native resolution, rendered at {1, ½, ¼, ⅛} without retraining. Isotropic / anisotropic rows differ only in footprint shape; the VKRayGS row re-implements only its paraxial z/f footprint sizing inside the same renderer.
| Dataset | Method | ×1 | ×2 | ×4 | Avg. | Corner ×4 |
|---|---|---|---|---|---|---|
| ScanNet++fisheye | 3DGEER (baseline) | 27.28/.944/.122 | 27.74/.918/.221 | 27.33/.895/.279 | 27.45/.919/.208 | 22.73/.871/– |
| 3DGEER (+ Mip-Splatting floor) | 29.49/.951/.111 | 27.19/.918/.216 | 27.25/.896/.272 | 27.98/.922/.200 | 22.40/.872/– | |
| 3DGEER (+ footprint floor) | 28.73/.949/.112 | 28.13/.921/.213 | 27.42/.897/.271 | 28.09/.922/.199 | 22.82/.879/– | |
| ConeGaussian (ours) | 28.80/.950/.111 | 28.32/.924/.204 | 27.52/.900/.264 | 28.21/.925/.193 | 22.88/.879/– | |
| Zip-NeRFfisheye | 3DGEER (baseline) | 25.87/.899/.150 | 24.92/.845/.278 | 24.13/.787/.409 | 24.97/.844/.279 | 21.72/.780/– |
| 3DGEER (+ Mip-Splatting floor) | 25.88/.899/.150 | 24.92/.845/.279 | 24.13/.787/.409 | 24.98/.844/.279 | 21.73/.779/– | |
| 3DGEER (+ footprint floor) | 25.91/.900/.150 | 24.94/.845/.279 | 24.14/.787/.409 | 25.00/.844/.279 | 21.72/.780/– | |
| ConeGaussian (ours) | 25.96/.901/.148 | 25.02/.851/.270 | 24.25/.795/.398 | 25.08/.849/.272 | 21.74/.781/– | |
| Mip-NeRF 360pinhole | 3DGEER (baseline) | 29.73/.898/.128 | 26.92/.790/.255 | 25.62/.720/.379 | 27.42/.803/.254 | 25.77/.718/– |
| 3DGEER (+ Mip-Splatting floor) | 29.93/.902/.125 | 27.03/.794/.252 | 25.74/.727/.372 | 27.57/.808/.250 | 25.94/.726/– | |
| 3DGEER (+ footprint floor) | 30.05/.906/.118 | 27.11/.798/.243 | 25.77/.728/.365 | 27.64/.811/.242 | 25.96/.726/– | |
| ConeGaussian (ours) | 30.15/.907/.117 | 27.18/.802/.239 | 25.79/.731/.359 | 27.71/.813/.238 | 25.96/.729/– |
Zoom-in: trained at the lowest resolution, rendered at ×1–×4. Corner ×4 is PSNR / SSIM over the normalized radial band r > 0.9 of the fisheye image.
| Setting | Method | 1 | ½ | ¼ | ⅛ | Avg. |
|---|---|---|---|---|---|---|
| FH → PHfull img | 3DGEER (baseline) | 27.42/.935/.257 | 29.99/.940/.203 | 29.66/.941/.156 | 28.34/.930/.136 | 28.85/.936/.188 |
| VKRayGS-insp. z/f dilation | 26.99/.933/.253 | 28.50/.937/.202 | 25.73/.919/.172 | 23.02/.871/.172 | 26.06/.915/.200 | |
| ConeGaussian (ours) | 27.38/.934/.255 | 30.14/.942/.199 | 30.28/.949/.135 | 29.56/.951/.083 | 29.34/.944/.168 | |
| PH → FHfull img | 3DGEER (baseline) | 27.01/.919/.234 | 27.47/.931/.185 | 27.53/.936/.139 | 26.66/.923/.120 | 27.17/.927/.169 |
| ConeGaussian (ours) | 27.01/.920/.230 | 27.47/.931/.181 | 27.69/.942/.124 | 27.02/.939/.077 | 27.30/.933/.153 | |
| PH → FHperipheral img | 3DGEER (baseline) | 23.82/.892/– | 24.11/.897/– | 24.32/.901/– | 24.12/.890/– | 24.09/.895/– |
| ConeGaussian (ours) | 23.83/.893/– | 24.15/.898/– | 24.44/.907/– | 24.36/.904/– | 24.20/.900/– |
Cross-camera rendering on ScanNet++ (FH = fisheye, PH = pinhole). FH→PH: trained on native fisheye views, rendered through the matched pinhole cameras without retraining; PH→FH is the reverse. Peripheral PSNR / SSIM over the radial band r > 0.7.
If you use ConeGaussian in your work, please cite our paper. The BibTeX below follows the public arXiv record.
@article{zhang2026conegaussian,
title = {ConeGaussian: Anti-Aliased Gaussian Ray-Tracing for Generic Central Cameras},
author = {Zhang, Deheng and Shi, Letian and Yang, Runyi and Li, Zhendong and Sun, Lei and
Wu, Kanzhi and Chhatkuli, Ajad and Paudel, Danda Pani and Van Gool, Luc},
journal = {arXiv preprint arXiv:2609.13397},
year = {2026},
url = {https://arxiv.org/abs/2609.13397}
}