This article is full of mistakes and misleading claims:<p>1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.<p>2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.<p>3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.<p>2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.<p>3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.<p><a href="https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/" rel="nofollow">https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...</a>
> 1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.<p>It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.<p>In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?
I specifically argue that it doesn't matter if it's $1 and base my argument and solution around that. So it's irrelevant where it is in 2035.
Even if it costs zero to create. How do You force people to pull from Your repo?
Basing your cryptographic advice on a 20 year old opinion-piece from somebody with no background in cryptography is not the flex you think it is.
Appealing to lack-of-authority without actually explaining in what way his argument is wrong is significantly worse.
But the Linus piece is sound<p>... for Linux<p>... and developers working for it constantly<p>the attack wouldn't work. Joe Schmoe? It's worse than just "being compromised"<p>You have repo of dependency locally, let's assume you downloaded good copy, the commits get compromised, you're safe.... right ?<p>Nope, if there is build server along the way and ESPECIALLY if it practices building from clean state every time, the build might be infected while your local copy is clean, giving no chance to notice it, unless your entire chain including local builds are reproductible AND you actually check it
The problem of a SH1 collision happening <i>by coincidence</i> is vanishingly low and theoretical.<p>Nothing else matters.<p>Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
Sorry, no.<p>When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.<p>Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?<p>Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?<p>Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
Dealing with potentially-hostile hosts is quite common, actually. See for example how most Linux mirrors work, or Subresource Integrity with HTML.<p>Turns out securing a service to transfer a single hash is a <i>lot</i> easier than securing a service to transfer gigabytes of data.<p>Even if I don't fully trust Github, it is still <i>incredibly</i> convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that <i>should</i> be secure.
I mean, Git commit signing should be used more often... then you can actually trust the person signing, not the distribution method<p>But you still need SHA256 for that
> Git hashes are not supposed to be a security mechanism<p>Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?
Because you're not killling two birds; you're not killing the security bird with a better content hash.<p>A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.<p>It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.<p>Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.<p>Nobody is going to fake your commit such that the fake has the same SH-1 hash <i>and</i> your GPG signature.<p>The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.
Please educate yourself what a merkle tree is. It's a well understood building block of various security systems, including certificate transparency (which explicitly uses sha256).<p>You refer to PGP signed Git objects, but you also argue:<p>> Git hashes are not supposed to be a security mechanism<p>Guess what the Git PGP signature is signing.
This is exactly right. A signature is only worth as much as the hash that it’s signing. And all the usual signature algorithms are signing a hash.
The GPG signature is not signing the git hash, if that's what you mean.<p>The GPG signature signs some kind of hash calculated over the commit, minus the GPG header, which is thereby added.<p>The git hash is then calculated over the whole thing. The git hash is on the outside, and not part of the signing.
> The GPG signature is not signing the git hash, if that's what you mean.<p>It kind of is - it’s signing the hash of the tree object, which is the actual thing that you’d attack with a hash collision
I understand that if we sign a commit with the help of some arbitrarily strong hash, it doesn't protect the parent commit(s). The integrity of the SHA-1 hash references to the parent commits is not in question, but the authenticity of those commits themselves.
No, not the abstract tree formed by a series of commits.<p>The actual git ‘tree’ object, which is the thing a commit actually points to, referenced by a hash in the commit. That <i>is</i> signed by the GPG signature.
That doesn't make a difference: with sha1 a malicious change in content will still result in the same content hash, so the signature will still be valid, and the commit hash will still be the same.
Only if the GPG signing process stupidly relies on the SHA-1 hash. I.e. if it takes an unsigned commit and signs only its SHA-1 hash and then creates a new commit with GPG headers. If that's how it works, that is massively stupid and can be fixed without forcing SHA-256 as a git hash. Just have the signing calculate its own digest for its own purposes.<p>That digest can be the SHA-256; since the infrastructure is there for it, signing should use SHA-256 regardless of what hash is used by the repository for identifying and linking content.
> Git hashes are not supposed to be a security mechanism.<p>Commit signing indicates otherwise.
One of my favorite fun facts about Fossil SCM (another source control by the devs of sqlite) is that they patched their use of SHA1 6 days after the shattered attack was published:<p>"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."<p><a href="https://fossil-scm.org/home/doc/trunk/www/hundredandone.md" rel="nofollow">https://fossil-scm.org/home/doc/trunk/www/hundredandone.md</a><p>To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.<p>That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
That's impressive. I suppose they had a more flexible architecture to make that change so fast.<p>Is there any writeup on why it was easy for them and not for git?
My guess is that it's less about the architecture and more about the blast radius and the number of users
Hmm probably nothing architecture wise. It's probably just the fact that fossil is developed by way fewer devs.<p>From the skim I read of this article it seems both projects arrived at the same solution: support both but make SHA-256 the default.
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.<p>Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
Also, interestingly, Git today does _not_ use a straight SHA1 because of these attacks. It uses `sha1dc`, a slower collision detecting variant that specifically checks for this vector of attacks. So currently, Git's SHA-1 variant is not susceptible to the SHAttered/Shambles attacks.
Agreed! This isn't a tech dig at all.<p>To me it's more of a reality of creating a tool with a huge active community and a community of contributors and creating a tool with a small team and small community.
It's easier to make world breaking changes when the world is really small. If git could magically just get everything and everyone to cut over and use git 3.0 in a magic instant, it wouldn't be having this problem.
Linus Torvalds in 2007:<p>> but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, Okay, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get. ... [1]<p>[1] <a href="https://www.youtube.com/watch?v=4XpnKHJAok8&t=56m20s" rel="nofollow">https://www.youtube.com/watch?v=4XpnKHJAok8&t=56m20s</a><p>So Torvalds used SHA-1 purely because he needed a hash function with no other property than identifying content.
For reference: <a href="https://git-scm.com/docs/hash-function-transition" rel="nofollow">https://git-scm.com/docs/hash-function-transition</a>
I don't understand why Git is not making the SHA-1 and SHA-256 modes far more compatible with each other.<p>SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.<p>But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date <i>or</i> referenced as part of the Jan 1 block.<p>And now it would be impossible to get a new SHA1 collision in to the repo.<p>The only new git features needed would be:<p>a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object<p>b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit<p>c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
Emily's talk does a pretty good job of summarizing the issues with intermixing the hashes: <a href="https://youtu.be/eJJp0RE7cd4" rel="nofollow">https://youtu.be/eJJp0RE7cd4</a>
I skimmed the video, and I didn't quite catch that. Near the end of the video, however, she did mention that interop is in the works[0].<p>In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.<p>Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.<p>[0] <a href="https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s" rel="nofollow">https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s</a>
Her entire section 2 (starting at 6:12) is basically about why git does not and will never allow mixing of SHA1 and SHA256.<p>The interop discussed is using copybara as a copy tool to move data from SHA1 based repos to SHA256 based repos and vice versa.
Thanks. I went back and re-watched that part, and her point's that "[a] tree's cryptographic strength is equal to the weakest hash algorithm anywhere in the tree," and that's fair. I also understand why Git 3.0 may not want to give user the choice, though I would still rather it be given.
> "[a] tree's cryptographic strength is equal to the weakest hash algorithm anywhere in the tree"<p>This is simply wrong IMO, for two reasons:<p>1. The attack on SHA-1 is a <i>collision</i> attack. Once you have frozen a hash, you cannot attack it with existing cryptanalysis. If there were a preimage attack it would be a different story.<p>2. Even if there were preimage attacks, one could freeze a mapping from SHA1 hash to SHA-256 hash.<p>In fact, #2 seems like en excellent design. Objects could reference such a mapping, and a repo could disallow conflicting mappings (the mappings would only be accepted if the mapped objects are reachable from the mapping and the mapping is correct).
They do give the user the choice, but the default is changing. My point is not necessarily to rip out the SHA-256 option, but simply to not make it the default. Because then people will create repos in that format that do not understand the ramifications, where the opposite should be true.
My point is not that it's the end of the world (or the end of Git), but that it will be painful and unclear and confusing to lots of people. That would be fine if it made a huge difference in trust or protection, but it's the wrong way to do that.
Hm but the date is stored inside of the commit. The only way we can know that a commit's date is authentic... is through its hash. If I can forge commits with any SHA1 hash at will, I can make a repository whose head commit has the same SHA1 as the one in torvalds:
/linux but where any commit was replaced by a malicious commit with the same SHA1 and a fake date. You have no way to detect that my repo is inauthentic other than through a deep history comparison. The whole idea behind a merkle tree is that just checking the hash of the top is sufficient to know the identity of the whole tree.<p>I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.
The date that a repo receives a commit is known to that repo. And a repo can stop accepting new SHA1 objects. And a SHA256 object could have a flag that says that no SHA1 objects may ever reference it.
The design of Git, as a Merkle tree, is meant to allow for use cases like this:<p>* I host a mirror of the Linux git repo.<p>* You download Linux from my mirror.<p>* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).<p>* You check whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.<p>This is a completely normal use of Git. People download from mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.<p>If I can forge my own commits to have any hash I want, this whole model breaks down. I can replace some old commit in the repo with my own forged commit with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.
> You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism<p>That's a 160 bit hash, which is SHA-1, which has the security properties of SHA-1.<p>Suppose you check out a commit with a given SHA-256 hash. That commit object represent the root of a tree where all the edges are hashes (and types, etc). I'm suggesting one of two designs:<p>a) (Simpler but weaker) If Linus has published that commit, then he is confident that he hasn't pulled in any too-new SHA-1 hashes and that there are no collisions present in what he thinks the tree is. So, by induction on the traversal depth, there is only one actual object identified by each edge, and those objects contain the hashes of their child edges, so those hashes are all correct.<p>This breaks if there is a malicious collision already in the tree.<p>b) (Stronger but higher overhead and more complex) There would be an object or objects, discoverable from the root by following only SHA-256 edges, that encode a duplicate-free mapping from SHA-1 hash to SHA-256 hash. The client finds and parses that and then, as it traverses the tree, each time it reads a SHA-1 hash, it computes the SHA-1 and SHA-256 hash of the referenced object, verifies that the pair is in the mapping and also verifies that the SHA-1 hash matches what the edge requires.<p>I think that (b) is genuinely cryptographically secure in the sense that, if you can construct a commit that has the same SHA-256 hash as an official upstream commit but different contents, then there is necessarily a SHA-256 collision.
For A), I don't understand what the point is? I never mentioned what Linus is confident about, I talked about what you can verify when you pull from my mirror. I could replace a commit from 2010 with a malicious one<p>For B), I would think this could work, but it's a completely different solution from what you proposed and what I responded to.
If there is one way enforcement (i.e. there is one point where the last SHA1 commit was signed by first SHA256 commit), I think it should be safe ?<p>The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered
This means, if you migrate your repo, every single commit message that contains text like: "please see commit <sha1>" will now be broken.<p>This will be a train wreck. I hope they don't release before adding compatibility modes to keep the existing sha1's around in the database.
Tools like git-filter-repo[1] support rewriting commit hashes in commit messages. git-filter-repo actually does it by default; see `--preserve-commit-hashes` in the manual[2].<p>[1]: <a href="https://github.com/newren/git-filter-repo" rel="nofollow">https://github.com/newren/git-filter-repo</a><p>[2]: <a href="https://htmlpreview.github.io/?https://github.com/newren/git-filter-repo/blob/docs/html/git-filter-repo.html" rel="nofollow">https://htmlpreview.github.io/?https://github.com/newren/git...</a>
I need git-filter-repo to rewrite entire documentation and also resurrect and rehire earlier employees to repeat their GPG signatures.
Sure, but that's not going to rewrite Slack messages, emails, GitHub links, docs
Came to say something exactly like this... The old sha1 handle needs to be still available in the same way an HTTP 301 redirect would work.
... do not migrate old repos? I'm not sure why people would do that. Or, if they do, why would they replace the current repo name instead of creating a different one and keeping the old one closed to make the references work.<p>I don't think this is going to be a problem at all.
Once this starts being actual pain, we will each vibe the replacement index creator (git already supports replacement objects), for back-forth conversion, populated on pack and object indexing.<p>For massive perf and mem use damage. But oh well. And then we will wait for official version
Since SHA-1 is already broken (just expensive in terms of GPU-time), then the text "please see commit <sha1>" is also already broken.
You can't attack an existing normal commit.<p>But also collisions there aren't a big deal. People will cite short hashes when referring to things and that's not "broken".
That is essentially only a second preimage problem, which is basically impossible.
There are plans to keep sha1s around in a database, but as far as I know, no way to transmit those, so they seem specific to individual forges. They can be recomputed, sure, but again, any signatures break and it's possible that in the case of an actual replacement, the recomputation is now wrong and not easily comparable. So what is the point?
From what I'd read, SHA256 in git is showing every sign of being another IPv6. In particular:<p>- It's implemented in a non-backwards-compatible way<p>- The benefits over the older model are a bit nebulous<p>- There's a large amount of tooling that needs to catch up, and little sign that there is movement there
The difference with IPv6 adoption is that the internet relies heavily on network effects: so long as some hosts only have an IPv4 address, you need an IPv4 address for full connectivity, but then if everyone has an IPv4 address anyway, there is no immediate need to migrate to IPv6.<p>(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)<p>This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.<p>If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
> Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting<p>Funnily enough, self hosting is why I can't use IPv6. I want vlan isolation, but only get a /64 from my ISP.<p>Fortunately the lack of IPv6 also isn't a meaningful loss anyway so whatever
I just started learning IPv6 with AWS since they charge $0.005/hr per IPv4. Maybe it will be more expensive in the future and eventually it will be the new default.
changing repos to the new IDs would break any existing links to content on the pre-migration repos.
> - The benefits over the older model are a bit nebulous<p>This is far from the case with IPv6!
There's a <i>massive</i> push right now from top down to have secure software supply chains. Google SBOM and SigStore. It's not an organic need but if you have government customers you don't have many options.
I thought this would be a snark but it's an extremely well put together argument against the "Hashmageddon".<p>If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.<p>I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
Technically, git's design (thanks to very smart people trying to solve this problem like brian and others) is _very_ easy to modify to different hashing algorithms now. A lot of amazing work has gone into this in recent years.<p>However, it's not a git problem. It's an ecosystem problem. It's that every git repo has to choose one and they're entirely incompatible with each other. That is the cost and the difficulty.
I think there is a question though when that will happen and if it will be in our lifetime. SHA-1 started showing weakness in 2005 (collision in 2^69 instead of expected 2^80. This was later brought down to 2^61 in 2011), the same year git was invented. Nobody has found a similar weakness in SHA-256 as of yet. SHA-256 is still at its design strength of 2^128<p>It took 20 years to go from vulnerability in sha-1 to having to replace it out of caution. There is no such vuln in sha-256 yet. It could easily be 25 years before we find one, and another 25 years before we have to do something about it. Perhaps longer. Will git still be used 50 years from now?
With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.<p>All it takes is just one collision to consider it broken right?<p>But hey maybe the attempt to fix it makes git controversial enough it falls out of favor, and nobody uses it anymore in 2 years, problem solved? sure.
> All it takes is just one collision to consider it broken right?<p>No, its considered broken before that stage. i.e. when someone discovers an attack that would allow someone to create a collision faster than they should while still being impractical.<p>> With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.<p>Computer power doesn't super matter, what matters is algorithmic breakthroughs. So far i dont think there are any examples of major breakthroughs of that type via AI, although perhaps i am just misinformed. Its still early in the AI revolution, it might still happen, but as it stands i don't think there is any reason to worry about that.
Perhaps not clear enough, but Scott was a cofounder of GitHub, so he knows a thing or two about git in the real world =)
So they’ve been talking about this for many years, planning, and finally announce when they’re going to switch the default.<p>So <i>this</i> is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?<p>I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.<p>This seems like a bunch of Monday morning quarterbacking.
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
Your argument is persuasive and well illustrated. I think the problem is the intro paragraphs come off as too certain of catastrophe which, when juxtaposed with your claim that "smarter people than me have been working on this", makes it sound like you don't actually believe they're smarter than you. The rest of your essay feels fair and not judgmental.
The plans have been on display in a cellar. Beware of the leopard.
A number of years ago when I heard about this, I was pretty angry and made a private fork of git immediately in which I tried to scrub away the SHA-256 bullshit. But that's basically just paddling upstream with a spoon for a oar.<p>The stewards of Git are going to do whatever they want, and there is nothing you can do about it if you don't have the clout to create a fork that takes the lead.<p>No amount of discussion will do anything because they've already decided that their view of the situation is correct. Git hashes are not just content identification but a digital certificate mechanism, and their collision resistance is a grave issue that must be fixed, the end.<p>You will be browbeaten in any discussion; it's not worth the energy in a world replete with issues.
>it will be an incomprehensibly expensive and ultimately valueless and avoidable global nightmare.<p>thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
Actually, I think sha-256 is possibly faster than the sha1dc variant that Git currently uses.<p>I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.<p><a href="https://lore.kernel.org/git/20260929112544.86511-1-scott@gitbutler.net/" rel="nofollow">https://lore.kernel.org/git/20260929112544.86511-1-scott@git...</a>
He means costly in terms of human effort and wasted time.
Core of the issue seems like github UX issues that they can solve
Yes every repo is either one or the other but you fix that by rehashing the entire repo. Everyone can do this independently. It's entirely possible to maintain to identical repos in SHA1 and SHA256 mode but for the most part I suspect once updated people will simply pull down the new repo and use git 3.0 as a required version.<p>As migrations go, it's reading as simple to me. You'll just have to backpoint the commit signatures. I must assume there's a backwards compatible reference for them in git 3, right?<p>Or drop them and reference the old structure in a dire pinch.
Re-hash the entire repo as in rewriting all history? Hooo boy will that be a mess, I deal with things which reverence commits by hash in repos <i>all the damn time</i>. There are thousands of them in every Yocto project!
Do you have any external references to any commits that matter, for example in your communication platforms (emails, Slack) or your bug tracker? Or, worse yet, in places where they aren't just text format references, but used for things like CI/CD caching decisions or security scans?<p>Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
"The migration to the new format is simple; Just re-write everything in the new format, but also keep the old format around forever too since data is lost in the new format!"
I positively don't care about the collision issue.<p>If you need to certify the authenticity of some code, and you've decided that a Git hash of any kind is going to be your certificate, you have a problem between keyboard and chair which is not fixable by stronger hashes in Git.<p>I don't want instability and churn in tooling.
Can someone more cyber-pilled than me explain what the actual risk with Git hashes being susceptible to collision attacks is? Obviously accidental collisions are problematic, but to my understanding the probability of that is still approximately zero.<p>Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a <i>far</i> from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
It seems people are missing the point: it's not even the submodule incompatibility that's going to become an issue majorly (like python2 --> python3 but worse), the main issue is the loss of traceability for repos that changes in place (which I assume many will do). Imagine what will happen to these:<p>- SLSA and Provenance or SBOM data in the supply chain security that uses commit hash. All the previous images are now pointing to a non-existing commit<p>- All the documentation and tooling as the article calls out<p>- All your traceability links from your project tool to your git repo, they will lose all the past data as it will be dead links<p>So I hope there IS NOT a migration path for in-place replacement!
Oh, agreed this sounds like a terrible migration path and shouldn't really be needed in the first place.<p>What I'm missing in the article is whether any Git server accepts replacing a SHA-1 identified object it already has. If it doesn't, then the distribution trust discussed holds, and keeping SHA-1 seems fine. Adding additional signatures seems fine for those who need transitive trust.
Ugh, I didn't know that SHA-1 submodules wouldn't be supported in SHA-256 repos. That changes the transition from painless to a major dumpster fire. Having to maintain converted forks, and use different hashes from upstream is going to be a mess.
Make <i>git init</i> use SHA-256 if git is invoked as <i>git3</i>, SHA1 if invoked as <i>git2</i>.<p>Plain <i>git init</i> could fail with a diagnostic: informing to use one of the two aliases or an option.<p>What people don't want is making git repos SHA-256 by accident and finding out later that they made repos not compatible with older git.
I don't agree that SHA-1 is much longer feasible for git.<p>But I also don't think that switching to SHA-256 must be painful. A git2->git3 converted repo could just store all the past hashes, so existing links don't break.
schacon: Really like the "Independent Tree Hash Headers" idea.
How difficult would this be to get this functionality into git?
Would it cause any breaking changes with older versions?
Have you discussed this with any git devs to see if they are open to adding it?
Actually, this entire blog post came out of a short chat at Git Merge a few weeks ago with Jeff King. I argued more or less this and he didn't _entirely_ disagree, though he has good counterarguments on the list over the last few years, so I don't really know how he thinks about it ultimately.<p>I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"
Also, functionally, this is incredibly easy to add to Git.
> We can go through years of this SHA-1 to SHA-256 migration and then quantum computers break 256 and we're back in the same stupid boat again.<p>SHA-256 is considered quantum safe by the NIST and is left out of PQC migration guidance entirely.
It was theoretical - the point was that maybe some paper is published or some new tech or issue comes up. Now we have to do this again. If we separate the concerns, then we don't have to deal with both as though they're one problem. We can deal with one thing for content addressing and another for trust and security.
Until it isn't.<p>Things like this have a tendency to to be revised as time passes.
For now. Turns out the pigeon hole principle still holds.
The post's argument that hash collisions are irrelevant in practice is not convincing at all. Basically they amount to:<p>1. Collisions aren't as bad as preimage attacks<p>2. Even if you made a file-with-malicious-hash, how would you get people to pull it?<p>3. Other attacks are a bigger problem (social engineering)<p>(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.<p>(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.<p>The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.<p>[1]: <a href="https://www.reddit.com/r/programming/comments/5vyhy2/webkit_just_killed_their_svn_repository_by_trying/" rel="nofollow">https://www.reddit.com/r/programming/comments/5vyhy2/webkit_...</a>
(2 counter) is impractical because all nodes of git will not replace objects if it thinks it already has it. So any attack has to assume this is the first time the node fetched, which is difficult before trust is established, which is difficult. This is part of the argument Linus originally outlined for this vector, which is that it only works for _very recent_ objects.<p>(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.<p>The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.
couldn't github reject a push that contains an existing hash in the repo?
It doesn't reject, but it will not replace. Same for a fetch/pull. That is another issue with this attack vector (that Linus also mentions) - it has to be the _first_ time that a node has seen this object. It makes the attack even more difficult than it already is (in like 4 different major ways)
I suspect everyone renaming their branch from master to main caused more unnecessary breakages and toil than this ever will.
The problem is existing repositories have SHA-1 commit hashes, and if you also end up changing those...<p>There's <i>lots</i> of tools that refer to git commit IDs. Some of those tools may even hardcode a commit ID to be 40 hex digits long. The fact that these tools are external also means that "oh, just rewrite the commit messages or code to refer to the new IDs" isn't feasible. The only way to not break the world is to let people refer to existing commits with their SHA-1 hashes in perpetuity, and it doesn't sound like git is set up to allow this in any way, which means that existing repositories have to stay SHA-1 in perpetuity and that will cause fun down the line if you start having to make SHA-1 and SHA-256 repositories.<p>Changing from master to main is a one-off change. It might require changing your scripts once to refer to 'origin/main' instead of 'origin/master', but other than that, there is essentially nothing more that needs to be done, there is no risk to historical artifacts that needs to be mitigated.
My prediction: 20 years from now everyone will still use SHA-1 git. That will be simply easier.
I run a small git/jj forge and for us it's already painful dealing with this. Can't imagine how GitHub is going to handle it.
A lot of replies here seem to be asserting that this "isn't that hard" without addressing the thing that makes it most hard: submodule compatibility and the breadth of tooling
The xz backdoor is basically his point in practice — that was a maintainer-trust compromise, not a hash collision.
I'm going to completely ignore the first part of this post because I'm not interested in arguing about how severe the issues with SHA-1 are. I think it's accepted that there are flaws.<p>So given that, I'm more interested in the arguments for why migrating to SHA-256 is problematic.<p>The biggest issue I see, after skimming over it, is the submodule breakage for new projects trying to link to old projects. This seems solvable frankly, but is the only serious issue I see. Everything else will be worked out as software is updated IMO.
The first reason the author lists for why this will be bad is only an "issue" on Git hosts that don't allow repo creation on push (which is brain dead of GitHub). Any other host, you push your new repo, and it will see the hashing algorithm, and receive the contents accordingly.<p>Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a <i>setting</i> which can be changed.<p>I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.<p>Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
I'm honestly still shocked any of this happened.<p>Prior to SHA1 we had MD5, a decade earlier. MD5 collision attacks had already been widely documented and known. It was the most obvious thing on Earth that this would happen to SHA1 too. Apparently, Linus never realized there was a need for cryptographic security and that the hash was purely internal.<p>Here's what I honestly think was a factor. I think C programmers fell in love with the implementation that you could throw around a fixed hash record on the stack. It's incredibly efficient. But it's an efficiency that doesn't really matter because as soon as you read from or write to a disk or a network or even memory, any cost saving is completely gone.<p>More than a decade ago, some people wrote a Java implementation of git (jgit?) and despite all their optimizations, it was (IIRC) only half as fast as C git. It is of course because Java at the time had no concept of stack values for non-primitive types so couldn't compete. Personally, I was impressed: only half the speed? That's pretty good.<p>For something that's only 20 years old, the Git SHA1 assumption is some of the worst technical debt we have in the modern era.<p>Here's another thought: when people make a lot of these programs, they often make the mistake of not separating the program version and the network protocol (or just the external API). So you end up with brittle client-server implementations where you have to upgrade both the client and the server at the same time because they lack a network abstraction.<p>The other end of the spectrum is video streaming where you have codex, container formats, transport protocols and so on.<p>What a mess.
they could just have taken the sha1 of the sha256.
Then `git fsck` wouldn't be able to tell if an object uses the sha1 scheme or the sha1∘sha256 scheme, so it may need to compute the both hashes to check an object's integrity. Also, an object name alone wouldn't indicate that the weak scheme shouldn't be used to check integrity, so a malicious sha1 object could be swapped in in place of an sha1∘sha256 object (if second-preimage is found).
> I pull it from there because I trust that GitHub has its authentication game together enough that it's unlikely that anyone malicious pushed something there without the maintainer's knowledge.<p>Hahahahaha, that's some level of delusion
Thanks for writing this. I'd only been loosely following it and I hadn't realised how bad this is going to be. I have repos with tens of submodules and it's going to be a nightmare if any of them switch to sha256 in place. Not to mention I won't be able to use any new projects unless I rebuild my repo and all the submodules therein.<p>I thought the master to main thing was bad enough but this is going to suck. And just like the master rename it achieves basically nothing.<p>What is it about these projects that attracts people who just want to change things for the sake of it? Real engineering means coming up with a solution for backwards compatibility. This is just irresponsible and, frankly, a fuck you to everyone who will be affected by this.
« people who just want to change things for the sake of it » — Not for the sake of it, but to “make the world a better place.” The intention is noble (well, mostly, at least let's assume it is). Of course, there is disregard for history and her deplorables (or not so deplorables), comes with being “progressive”. Which is why I like Windows better (not 11 though, will probably have to go back to Linux at some point).<p>« What is it about these projects » — Maybe that they're “at the forefront.”
This seems like Y2K fud.<p>The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.<p>This is very very easy to fix if you run into it.<p>1. Adopt git 3.0 if you can with sha256.<p>2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.<p>Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
> sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue<p>If you read the OP article, the entire point he's making is that this would never happen, because a hash algorithm being "broken" doesn't matter in practice, because true supply chain security has nothing to do with file hashes.
It's easy for _one person_ to fix. It's not easy for the entire git ecosystem as a whole. GitHub, large internal corporate git repos, CI/CD systems, projects with submodules, etc. The second half the article explains all of this.
It’s already broken, but even though it’s broken it’s hard to generate git collisions because of the repo metadata. It’s easy to generate (for instance) standalone PDFs with identical hashes, but doing this with git in a useful way is much harder.<p>That said, it’s still a good idea to migrate to a more robust hashing algorithm. Defense in depth, etc. Just because it’s a difficult migration doesn’t mean it shouldn’t be done.
You can certainly do this, as I said, this is Google's backup plan. But defaults matter. People will start running this and getting repos that are uselessly incompatible with other repos, tools, libraries and server instances. Having it as an option is one thing. Making it a default will cause a lot of pain for people who don't want to care about this.
> Adopt git 3.0 if you can with sha256.<p>Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
For the record, Y2K was not fud. It was very real, in a long list of datetime problems that are to come. Further datetime problems are coming at scheduled dates.
It was definitely FUD. There was a real problem (date counters would roll over), but the impacts of it were so ridiculously overstated that it eclipsed any sane discussion of the issue. We had people at the time predicting that planes would literally fall out of the sky when Y2k hit, which was never a realistic possibility.
> which was never a realistic possibility.<p>Because a lot of work was done to prepare and fix potential issues.
That's the problem with deniers. When responsible persons take preemptive action to prevent tragedy, like with Y2K, the diners say it was FUD. When people don't take action, like with climate change, they say it wasn't important considering it's not them who's dead, totally discounting those who have suffered or died as a consequence. In summary, the deniers are so incompetent that they can't be trusted to correctly maintain a car, let alone civilization, considering they would never even the replace the necessary parts at the right schedules in their car.
[dead]
…
Would the author feel the same if git had used MD5 instead of SHA-1?
I do actually literally write in this that if it was MD5 it also would not be a problem.
They address this very theoretical. In short: Yes. Which makes sense if you don't treat the hash as a form of security against malice, especially in the case of attacks that are already impractical, which is the entire thrust of the article.
The hash isn't the security, the distribution is.<p><a href="https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/" rel="nofollow">https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...</a><p>As linked by another commenter in this thread, Linus worked out years ago that even if someone inserted a malicious object into the kernel repo, it would at best be a nuisance and not a major concern.