I'm rather surprised that it doesn't take the z-buffer as an input. I would have thought that would have provided useful information, it's one of the more useful forms of contolnet.
The official one seems to do, as well as other info from the engine (I think remember their mentioning LOD/UV map hints in one of the public demos, or articles, a few months back--or it might have been an Unreal Engine podcast)
The technical report suggests it does not use the depth buffer: <a href="https://research.nvidia.com/labs/adlr/files/DLSS5_Report.pdf" rel="nofollow">https://research.nvidia.com/labs/adlr/files/DLSS5_Report.pdf</a><p>> The inference interface uses the engine-rendered RGB image as a dense, registered observation of visible scene appearance. It provides dense, pixel-aligned evidence for object support, occlusion boundaries, composition, and local material properties; engine motion vectors separately provide temporal correspondence.<p>> Existing image generative models commonly rely on text embeddings, exemplar images, or spatial control fields <i>such as depth</i>, edges, segmentation, and pose [...] These conditions are effective for general-purpose generation and editing, but they do not uniquely determine the object identities, materials, visibility relationships, lighting decisions, and pixel-aligned detail contained in an engine-rendered frame. DLSS 5 is therefore conditioned on the rendered frame itself.
Obviously they tried it multiple ways have brought receipts, but nonetheless it seems surprising that it wouldn't be of benefit to bring as much of that kind of metadata to the model as possible. You'd think depth and segmentation in particular would basically just be a straight shortcut without which the model spends its own time and effort re-deriving that stuff.<p>I'd also be interested in how post-processing fits in with this. Like if you've got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.
I think it has more to do with what kind of data they have access to at runtime - IIRC DLSS upscaling has only required the previous frames and motion vectors, so requiring depth buffers would mean it was no longer a "drop-in" replacement<p>> I'd also be interested in how post-processing fits in with this.<p>I <i>think</i> screenspace effects like film grain and tonemapping are excluded in the same way UI elements are rendered separately from the game.
I think this is mainly so it can use the existing hooks for DLSS upscaling without requiring changes to the renderer, AMD is working on a comparable method which uses adapter networks to slot normals and material properties from the renderer into the diffusion model: <a href="https://gpuopen.com/learn/temporally-stable-generative-illumination/" rel="nofollow">https://gpuopen.com/learn/temporally-stable-generative-illum...</a>
Even relatively small RGB -> depth models are pretty good. Which kind of implies depth is well encoded in the RGB, and adding depth would not really reduce entropy, while costing bandwidth.
> bit-exact against the original<p>What kind of sorcery is this ? Very impressive work !
Well it's doing the same math as the original, apparently. Hard to do but makes enough sense.<p>With LLMs you can do whatever you want pretty much. I have upstream CUDA running llama.cpp under unmodified Nouveau on one of my boxes. Why? Well, why not?<p>I also have a modified Nouveau driver that, with the help of more and newer blobs, gets reclocking working for at least most of Pascal/GTX 10 series. I would love to try to upstream it but it desperately needs to be rewritten with that intent. Too much ugly garbage. Still, I wanted to know how possible it is. Possible, it turns out. Modern LLMs can blackbox analyze the real driver quite well, and debug the Falcons themselves. It's very interesting. People say coding is dead; I think it's probably not really true. However, it is certainly changing. I think someone less skilled than me could beat me to the punch with enough determination. That is interesting.
Given that you bring the weights from NVIDIA's DLSS. So basically, the repo contains reverse engineered machinery that produces the exact same output given the same model.
Claim != Reality most of the time
If done by a human, yes.<p>Nowadays it takes one well written prompt to a frontier LLM to produce something like this.
That would be still impressive, even though more distributed among the AI builders and all those nameless code contributers etc.
So… that kind of sorcery, I guess.
The load-bearing kind.<p>LLMs are really good at deobfuscating or even decompiling code.
Jensen : Nobody needs to code anymore...<p>Programmer: OpenDLSS...<p>Jensen : Wait. Not like that! (╯°□°)╯︵┻━┻
You think Jensen even thinks about the gaming market anymore?
Hope it will get a Linux port soon
But I thought Nvidia “loves open source” (they don’t) and they are now a supporter for open source and open weight models by acquiring Huggingface? (They don’t actually care)<p>But the line is drawn when it involves CUDA and any part of their closed source compilers (nvcc).<p>There are obvious reasons why they are closed source, but it’s becoming pointless since Deepseek have open sourced their AI compiler and compute libraries with DeepGEMM and eventually they will catch up.
They only care because open-weight helps to drive the GPU business notably thanks to inference providers (Baseten, Together, Mistral, ...).<p>At least their support helps the open-weight ecosystem.
Depends on which open source you are talking about, like every single company contributing to FOSS.<p>They care when the agendas align, and they don't when they won't.
They certainly do _care_. Embrace, extend, extinguish.
They don't. I got threatened with a lawsuit after I suggested (submitted patches) they fix some of their buggy kernel code.
Almost 8ms on 1080p resolution seems extremely expensive, does the original also eat into the rendering budget as much?
Yes, see the numbers below. In most games, it's basically unusable if you want to play on 60 FPS or above unless you have a 5090.<p>The current implementation is more of a tech demo than a practical way to play games (+ officially it's available in 1 game). It's _fast enough_ to make some impressive YouTube videos but you most likely won't want to play anything with it yet.<p>Nvidia has stated that they're still working on improving the performance. No doubt future hardware generations will also include further hardware optimisations.<p>The potential for this kind of technology is pretty awesome, especially given that people have also found ways to add this to emulators.
I was able to run it in Skyrim on my 5060ti at about 60fps. First game that almost used all 16gb of vram
I really hope they do, but the market for gaming cards is looking mighty bleak right about now. 5090 is up over 80% since November last I checked and my 5080 is up 50%.
NVIDIA removing all mentions of gaming in their financials doesn’t bode well either, and from a fiduciary standpoint it would be negligent to sacrifice any capacity for higher-margin AI chips to make gaming cards.<p>Again, I hope I’m wrong and we see new cards summer/autumn 2027, but I would not bet my savings on it.
It does OK on most of the 5 series cards, but you can only do 4k on the 5090 basically. And there are lots and lots of knobs to tune and experiment with.<p>In some cases, it seems that lowering resolution and graphics actually produces better DLSS5 output (but it varies)<p>It's <i>really</i> good at making older gen games look remade/remastered.
> from a fiduciary standpoint it would be negligent to sacrifice any capacity for higher-margin AI chips to make gaming cards<p>Corporations do not have a fiduciary duty to seek maximal profits. This is a myth.<p>They are given wide latitude to decide what's in the best interests of the shareholders. Keeping a less-profitable offering alive just in case the current big offering doesn't pan out in the long term would easily be defensible in court.<p>It wouldn't even be a challenge. Courts are loathe to question the judgment of directors and executives. The reasoning is obvious: why in the world would a judge have better knowledge of how to run a company than the people whose jobs are to run the company?
You are correct, I should’ve used financial not fiduciary. They make more money selling AI hardware than gaming and it’s a zero-risk transfer because if AI collapses gamers will not “vote with their wallets” and not buy a new card from NVIDIA, and they know it.
And I don’t even disagree with them here, if you can make 10-100x more money doing less work, why wouldn’t you?
Modders have added the ability to use the upscaler after DLSS 5, personally I don't know how sound this method is but the quality is pretty good, and allows to play games with DLSS 5 at 4k 60FPS with something that is not a 5090.
Worth remembering it is running a single-step diffusion model working in pixel space to generate each frame, it's a technical feat in itself that people are even using the words "frames per second"
Yes the original is very expensive. It depends on the resolution and the GPU used:<p>RTX 5060: 9.9 ms at 1080p<p>RTX 5070: 10.2 ms at 1440p<p>RTX 5080: 13.7 ms at 2160p<p>RTX 5090: 8.2 ms at 2160p<p>Source: <a href="https://www.youtube.com/watch?v=3EfLjmdG29Q&t=600" rel="nofollow">https://www.youtube.com/watch?v=3EfLjmdG29Q&t=600</a>
The original does tend to reduce the FPS by half or more
probably, going by the reported massive performance hits
> A Vulkan reimplementation of NVIDIA's DLSS 5 Neural Rendering network, bit-exact against the original.<p>Bit-identical, I swear I heard that somewhere before.
At what point does this neural rendering take away the human touch on the art styles?
How useful is this without weights?<p>Isn't the mote that Nvidia has is they work with studios to generate the training data from the game, then they ship a model per game?<p>Or is my knowledge outdated here and they're just using a single generalised model?
I assume this is meant to run with the weights people extracted from the latest NBA game, where it was first trialled.<p>> Isn't the mote that Nvidia has is they work with studios to generate the training data from the game, then they ship a model per game?<p>That was true for the very first version of DLSS, from DLSS 2 on the models have been universal - the per-game adjustments are done on the inference end by changing the effect intensity or masking out objects<p>They have a technical report on the neural rendering part of DLSS 5 which goes into it: <a href="https://research.nvidia.com/labs/adlr/files/DLSS5_Report.pdf" rel="nofollow">https://research.nvidia.com/labs/adlr/files/DLSS5_Report.pdf</a>
I don't really understand why the weights aren't included... US law says they can't have a copyright, no? Maybe they're concerned about other countries or cautious about a litigious Nvidia.
They used to ship one model per game but now there is a single model, however they still do minor updates to it presumably to fine-tune it on new games
>they ship a model per game?<p>They don't. Only DLSS 1 was trained specifically per each game.
> they ship a model per game?<p>there's no way that's true!?
It used to be in dlss1. I think it's completely been put to pasture now though, it's way too much work and can't really cover some of the main things people actually want to use dlss5 on, for instance Morrowind.
its not, from what I've heard the difference in quality was not worth it
What kind of dataset is used to train DLSS 5? Do they need to generate synthetic image pairs first?
I think it's interesting that Nvidia is so interested in producing the hardware that fuels the future of software development, given that their primary business advantage is their software moat. This is an interesting project for sure, but turning a bunch of GPUs at Zluda[0] (an open implementation of Cuda) could be far more destructive for them, right?<p>[0] <a href="https://github.com/vosen/ZLUDA" rel="nofollow">https://github.com/vosen/ZLUDA</a>
Am I the only one who feels a sense of disinterest in a project where the main README is LLM-generated? Does the author not have time to write what they did and how it's used?
I'm more upset about it being factually wrong, e.g. both mentions of "git-ignored" are absurd (why would you mention it if it's not in the repo?) and wrong (they are in the repo).
I notice this, that AI likes to write about things that are not in there. Like i review AI generated output, notice unnecessary things, and asks AI to remove that. So AI removes that and adds that "this and that, that was used or described like this, was removed because bla bla bla" to the document.<p>I think its somehow needs to talk (write) about the things that are in the context and removal is there so AI predicts that it should be there.
Yeah, I call it bugfix storytelling. Once upon a time this class far far away had this red hooded method...<p>Especially egregious if both adding and removing the thing happens in one commit. Git should be telling the story, and if it can't then there _is_ no story!
I spend so much time cleaning up AI comments in the codebases I work on, it's maddening. I could have an instruction to not allow it to write comments at all, but some of them <i>are</i> useful.
Yes it's constant. And in any article it's full of "The X does Y. The Z does A. But the B stays silent."
AI writing is just bad, this things is really noticable - but even in READMEs they look superficially OK until you read them.<p>AI probably should not be writing docs, commit logs or comments.
If you think about it, actually the author did write what they did (nothing), and also how it’s used (it isn’t).
Exactly. The code, whatever, it's for machines so I don't really care if it's by machines as long as it works. But if you can't even be bothered to think about the human-facing parts of your thing like docs and UX, I'm not really interested. It just feels cheap (in the bad sense) and offputting.
Agree.<p>I do feel like there's merit to having an open source implementation of anything, no matter who/what wrote it. I'm just hoping the results are validated well.
> what they did<p>bold of you to assume the code wasn't llm generated as well.
you think a human wrote the code? in a month? Do you think that's air you're breathing? might be time to challenge preconceptions
I'm fine with it
If the README is >90% AI generated and it is as long as a novel, I am not going to read it and will assume that the author did not read or write it either.<p>Unfortunately it is slop, beyond the comprehension of the author unless they are experienced with DLSS internals to explain it in depth.
"Am I the only one who feels a sense of disinterest in a project where the code is LLM-generated? Does the author not have time to code the project?"<p>This is how I feel about every single project announcement on HN recently, they are already bragging about models all over the place, why shouldn't they go full way down being replaced by the Borg?
Slop is a new language and you will learn it read it
So that means AMD implementation is on the horizon?
Something will this would historically guarantee a Senior Staff+ position at Nvidia. Wondering why Jensen doesn't put money where his mouth is ("were seeking exceptional engineers blabla") and offer him a job?
> It takes one rendered frame (a low dynamic range proxy of it, three lanes of Gaussian noise, the previous frame's output reprojected, and five conditioning scalars) and produces four f32 channels per pixel: an RGB residual and one temporal-blend logit.<p>> The temporal path is implemented, but in the demo: the network's history input lanes and its per-pixel blend logit drive a reprojected feedback loop (docs/frame.md). The dlss5vk tool runs single frames with no history, which is what the reference captures were made with.<p>From this I assume the network uses the (via motion vectors) reprojected previous frame in order to increase temporal stability, i.e. similarity over adjacent frames. But this isn't strictly necessary, and apart from it, DLSS 5 is a pure post-process filter. So you could apply it to an old animated CGI movie like Final Fantasy (2001) [1]. Which should make it look significantly more realistic, at the cost of some flicker or other temporal instability.<p>One could also apply it to still images, like old renders from Tomb Raider [2], where temporal stability is not a factor. The difference to conventional text-to-image models with a "make it photorealistic" prompt would be that DLSS 5 strongly adheres to the underlying geometry.<p>1: <a href="https://www.imdb.com/title/tt0173840/" rel="nofollow">https://www.imdb.com/title/tt0173840/</a><p>2: <a href="https://www.tombraiderchronicles.com/images/artwork-high-resolution-tomb-raider-4/tomb-raider-4-artwork-high-resolution-71.jpg" rel="nofollow">https://www.tombraiderchronicles.com/images/artwork-high-res...</a>
Seems quite similar to this repository?<p><a href="https://github.com/aloshdenny/open-dlss" rel="nofollow">https://github.com/aloshdenny/open-dlss</a>
The maanHimself repository appears to be 35 hours older than the aloshdenny one based on `created_at` from the GitHub API. maanHimself's `pushed_at` predates aloshdenny's `created_at` too.<p>That doesn't guarantee maanHimself is the original author, but it's looking likely.
Same exact commits at the same time as well, but different repository names and authors. What the hell?
bit-exact against the original )<p>just unsure who's original ))
[flagged]