This article is full of mistakes and misleading claims:
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.
3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.
I say it's impractical to exploit, which I think everyone agrees with.
Impractical for an individual, definitely. For a large org, maybe, but if the payoff was big enough? For a nation state level actor intent on doing something, absolutely not.
The go-to example is Stuxnet. Some countries wanted to attack Iran's nuclear enrichment programme, so they spent 5 years developing a worm that used multiple zero day exploits to attack a specific controller in a specific model of gas centrifuge. Could Mythos write Stuxnet? Unlikely, but a knowledgable team with access to it could probably write it in a lot less than 5 years.
'impractical' has very different values for different groups.
Yes, and in the days of major supply chain attacks and state-sponsored near catastrophes like xz, "it's probably good enough" starts to look incredibly naive.
Linus's "what matters is distribution" comment also doesn't make sense when merge effectively is distribute. Which, again, is the reality of supply chain.
Okay, fair enough.
But does that make sense as a default setting then?
I can see that some things might have a risk profile that might possibly make all this costs still worth it, but does it make sense to have these unicorn projects effectively blow up 20 years of ecosystem?
Shouldn't the extra cost of doing something out of the ordinary be carried by whoever does something out of the ordinary?
This feels like a bridge to be crossed when one gets there (if at all).
__
FWIW, we actually do have a choice here. No one is forcing the industry at large to adopt an unpatched git 3.0 binary built from a source that makes that a default.
This should be a trivial overlay to carry around with effectively no downsides. So convincing whoever is steering that ship doesn't necessarily matter, as long as enough sane pragmatics agree on how defaults should actually be.
We are migrating all of PKI to the more costly and less efficient postquantum cryptography, even though nobody will reasonably use a quantum computer to snoop on your home IoT daily reports. I mean, I assume that what you are doing on your free time is not worth governmental attention.
The rationale of mass migration is that if you don't impose it, nobody migrates. This has notably been the case with famously insecure SSL parameters (512 bits RSA keys, PKCSv1.5...). And many companies may believe they are not critical, which might be true until it is not.
Case in point: you manufacture walkie talkies and suddenly your products have bombs inside. Or you maintain a compression library for free and suddenly you are shipping a backdoor to all Linux products.
Stuxnet is essentially "boutique malware". You can buy it/have it built with enough money/resources.
Weaponizing a cryptographic algorithm with some theoretical vulnerabilities (but no by-design backdoor built-in) is a totally different game. And TFA is right that it's a dumb endeavour. You can probably "stuxnet" your way in for much cheaper.
> Impractical for an individual, definitely. For a large org, maybe, but if the payoff was big enough? For a nation state level actor intent on doing something, absolutely not.
In all the years since 2017, with all the orgs having huge GPU-filled data centers (and an interest in software security) has anyone demonstrated a real git collision?
Some systems have other properties that mitigate or prevent second preimage attacks - for example when you get an SSL certificate, CAs randomise the serial number. So an attacker can’t choose the checksum of the data the CA signs. Perhaps something in the design of git is similar?
And it’s also the question of different use cases.
Asking to trust in an authority (while the main authority Microsoft/GitHub has is essentially figuring out enterprise sales well enough to be acquired by a company desperately needing developers after fumbling badly in the 2000s) is exactly the opposite of my stance - it is a large corporation, with heavy employee rotation, with substantial exposure to various forms of regulatory pressure and to various forms of corruption.
Which is why the cryptography exists to prove the developer-to-consumer trust without trusting the intermediaries. Yes, I do check GPG signatures. Yes, I do include a git commit hash in the binaries I build. And surely I want to make sure that this doesn’t mutate because some unknown engineer at GitHub had a bad case of gambling debt.
Impractical for everyone. Because if you want to do this, there are better attack vectors, even if second preimage was easy, even though it is impossible.
> 1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.
In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?
The XZ incident would have been so much worse if it would have involved a collision, pushing one object variant to github.com, and one variant to git.tukaani.org.
Then you would have security researchers making conflicting claims depending on which repository they first pulled from, even though they are on the same git commit hash.
In GitHub, fork networks are represented as one repo on disk (this is the main optimization that makes community pull requests possible at all). So you could fork a repo and then push an object with a hash collision to it to change a file in the original, in theory.
(In practice this is harder, as the article mentions, because the new forged object would have to be a valid gzipped git object of the same length. And GitHub probably knows about this type of attack and might just, for example, prevent existing objects from being overwritten)
the attack wouldn't work. Joe Schmoe? It's worse than just "being compromised"
You have repo of dependency locally, let's assume you downloaded good copy, the commits get compromised, you're safe.... right ?
Nope, if there is build server along the way and ESPECIALLY if it practices building from clean state every time, the build might be infected while your local copy is clean, giving no chance to notice it, unless your entire chain including local builds are reproductible AND you actually check it
In agreement that this is good old fashioned cargo-cult security theatre, but tom7 also coined a more catchy phrase for this, he calls it "toxic max-security." http://tom7.org/httpv/httpv.pdf
Conveniently Tom didn’t mention anything about Edward Snowden and what he published. That was basically start of TLS everywhere.
Then he didn’t mention ISP idiots that were actually injecting ads to cute websites like Tom’s. I hope Tom likes when his website is used by ISP to make money on ads he doesn’t have any control over.
Then he goes on to criticize certificate transparency, but it works. Companies got kicked out from trusted root program because they were doing stupid stuff like making certs they shouldn’t.
Let’s not forget glorious state of Kazakhstan where without TLS they would just listen to all traffic - well with TLS they were trying to pull MITM but were uncovered and got their stuff removed by TLS ecosystem.
If I recall, tom7 was at odds with chrome throwing up a warning at users trying to visit his website because he didn't support TLS on it. He wasn't against TLS. Calling them a TLS antivaxxer is not accurate.
You are basically resolving a non-existent security problem by generating a far bigger security problem, because I'm 100% sure that a ton of software just assumes that a git commit hash fits in a `char[40]` and thus will buffer overflow like hell if they try to operate on new repositories.
And we are talking about who knows how many tools that work with git built in the years, and this is also made it worse from the fact that most tools just invoke the git binary and capture its output instead of passing from a library.
I like more the solution proposed at the end of the article, do not change sha-1 but instead, if you are relying on git commit for security purposes (that was never the intended use) add another header to the git object with a sha-256, so that with the small expense of computing the hash twice you don't break 20 years of existing tools that make the assumption of the git commit being 40 character long.
I would have liked sha-256 truncated to 40 characters (by analogy to sha-512/256 that'd be sha-256/160). Cryptographically that's probably fine. But 160 bits is not a lot. And probably fine doesn't tend to inspire confidence in the field of cryptography. It's not a very well studied scheme
An advantage of a hash with a different length is that a full length sha1 commit hash and a full-length sha256 commit hash can't be confused for each other
Persistent problems with increased hash lengths seem unlikely: if bad git clients have erroneous truncations or buffer overflows, they can either fix them instantly, problem solved (they have had many years to implement and test SHA-256), or fail to fix them and be written off as obsolete and incompatible: problem equally solved.
The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.
Nothing else matters.
Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.
Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?
Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
Dealing with potentially-hostile hosts is quite common, actually. See for example how most Linux mirrors work, or Subresource Integrity with HTML.
Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.
Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.
Rumor has it that GitHub has a flat namespace for commits. They don't store "user1/repo1/abcd1234" in one file and "user2/repo2/abcd1234" in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes because the odds are infinitesimal that it'd ever matter, than voila, I've updated your repo by writing to my own.
Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.
I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.
> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?
Yes, they ought to be good enough. That has always been git's security model.
Note that the commit doesn't have to be communicated over the same channel as the git data.
> Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Then they're vulnerable. But what about somebody learning about the trusted hash in another way, e.g. a build server getting an internal call authenticated by an authorized developer?
Just because you can think of an insecure way to use git hashes doesn't mean there aren't any other, secure ones.
> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?
How does this matter? When I have a machine that I trust and a git hash that I trust, I don't need to rely on the transport mechanism. As long as the content cannot be forged to match the hash, the transport is completely irrelevant.
> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries
Wait, why is anyone expecting that to be a good idea at all?
Sorry, no. For trusting code there is code signing. A sha-1 hash is cryptographically the wrong approach for it or we wouldn't have RSA and ECDSA algos.
>I expect the code to be exactly what has been reviewed by me under that hash
Sorry, no.
Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
> Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
What? Uncircumventable? Logically equivalent, I read your statement as "So if all cars are not blue then they must be red"? This does not follow...
I like the last part of the article, which proposes a (very reasonable sounding) extension for people who care (more) about their code-sec. You should be able to swap out your hash algo without having to rebuild your content addressing system.
> Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
This doesn't seem to be a useful definition. Would you classify every computable algorithm as insecure, because by generating a random bitstring, there is a (very, very low) probability of guessing the hash/secret key/solution/signature?
> Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
Cool, you've just defined the foundation of signatures and web encryption as insecure. What next?
Also, you can only be so certain about any piece of data no matter what you do. With a non-broken hash you can make the collision chance be a trillion times lower than the chance you're hashing the wrong data to begin with. That's as good as gold, well actually better than gold.
Because you're not killling two birds; you're not killing the security bird with a better content hash.
A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.
It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.
Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.
Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.
The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.
The attack is I pre-author `Makefile => foo: echo "hello"; bar: echo "world"` along with `Makefile => foo: echo "hello"; bar: rm -rf / ; /* $ELDRITCH_SHA1_SPIRITS_GO_HERE */` that both hash to `ff1234...`
I then prepopulate the repo with `echo "hello"`, wait 6-9 months, then submit a commit for `echo "hello" ; echo "world"` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.
I then have free choice as to whether I serve them "hello world" or "hello && rm -rf", and THAT's the plausible problem to avoid: the ability to "cloak" content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU's.
You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default/lazy/strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... => sha256:abc123456789...` "transparently" rather than the horrific situation you're laying out (correctly!) that forks the ecosystem in to "longhash" and "shorthash" when most repos don't even care in the end.
Please educate yourself what a merkle tree is. It's a well understood building block of various security systems, including certificate transparency (which explicitly uses sha256).
You refer to PGP signed Git objects, but you also argue:
> Git hashes are not supposed to be a security mechanism
You do understand that most code signing is just PKI over the top of SHA hashes, right?
The only substantial difference is that the author vouched for these particular snapshots of code, in a way where nobody else can intercept the communication and substitute a completely different hash than one the author previously signed.
In fact if you can find a SHA collision, you can peal the signature off of the legitimate payload and slap it on the colliding one.
Exactly! The recommended way to use GitHub actions is to use their deploy hash as the version to prevent hijackings. It was never the intention (nor would this have been envisioned when git was created), but security needs to be applied to the way tech is used.
Package managers even use it. E.g. you can have a cargo dependency pointing to GitHub at a specific commit. It's definitely intended to provide end to end security without depending on GitHub being secure.
> Git hashes are not supposed to be a security mechanism.
They absolutely are, both when you're using signed commits (i.e. GPG, SSH, and S/MIME) or just identifying a given commit/repository state by hash via a secure channel and then serving object data over untrusted transports.
It's entirely possible that you don't use either, but that's certainly not true for everybody.
Calling anyone using either to be "doing it wrong" is borderline gaslighting: Git used to have these security guarantees, and just because they're now broken doesn't mean they were never there in the first place, or that it was stupid to rely on them.
That's what a lot of the responses are missing, which is why the response to them in turn is both no and yes. There's no published threat model that I know of for the properties that the hashes are supposed to be providing for git, which means anyone can make up any required property they like and then confidently state that SHA-1 provides or does not provide it, see this discussion for examples. The result is, to use another Linus quote as the OP has already quoted him in the writeup, "people wanking around with their opinions".
We use SHA-1 in our storage mechanism, which predates git. There is a (quite long) written threat model. Someone being able to generate collisions with an enormous amount of effort under just the right conditions is not a threat under that model.
No, that argument misses the point. Cryptographic hashes enable you to trust, that I only have to review the changes since the last trusted commit. That is more important for the developers and maintainers of the software and less for the users.
In order to work for the users, a thing first has to work for its makers.
There are a lot of tools where the lines blur between “for the users” and “for the development team” because the users benefit from some things that make the developers’ lives easier.
You could pull a malicious colliding PR and reject it. Then you pull the other half of the collision without realising it is, and it's something good and you merge it. But your CI server already saw the malicious one and thinks it's the same, so builds the malicious code
The important point is that the switch is breaking backwards compatibility. The proposed solution, a new independent hash just for verification makes sense.
This is the top comment and yet says absolutely nothing to refute anything in the article, merely declaring that it's "full of mistakes". I feel like people haven't read the article, or this comment, and are essentially religiously predisposed to so-called progress, no matter the cost.
Your points are full of mistakes and misleading claims.
I don't simply mean to be disparaging - its important to security that the people making the decisions are a) competent b) can read, otherwise any "security" decisions they are making are at best probably insecure, and at worst, causing active harm and insecurity, DOS, etc...
Generally, I think you didn't read (or at least comprehend) the article:
1) Your assertion is factually incorrect. Practical proof of concepts are not "CVEs exploited in the wild". The article is not claiming that SHAttered is not correct, it in fact references it.
2) I don't think you read the article. The article claims the exact opposite, and in fact addresses the issue WRT to distribution.
3) Your sentence here is very confusing - I'm going to give you the benefit of the doubt and presume that you mean to say that the problem here is in the security of the naming. HTTPS itself has nothing to do with distribution security, that would be DNS/SecDNS (IFF you are using git+http protocol, then HTTPS is relevant, but not to git otherwise). But this is exactly what the article was talking about, the distribution is the security, not the hashing algorithm.
Notably, a few of the featured author's reservations appear to be addressed. According to the Git docs:
- Objects can be referred to by their old, SHA-1 name or their new, SHA-256 name. This means old refs in docs and comments and such remain valid. The mapping between SHA-1 representations and SHA-256 representations appears to be intentionally bijective a.k.a. 1-to-1 (assuming no hash collisions), so that it could be re-computed on demand. The constraint of bijectivity appears to be the source of some limitations, ex. no mixed repos and submodules needing to match hash algroithm, but also bijectivity has strong benefits like the following items.
- A bi-directional dictionary is maitained from SHA-1 to SHA-256 names so translations between the two don't required re-hashing objects. This table could be recomputed on demand due to the bijection between names; it's only a performance optimization.
- A local SHA-256 converted repo (including an SHA-256 converted submodule) can interoperate with an SHA-1 only remote transparently to the remote server by translating names using the lookup table.
- SHA-1 based GPG signatures will be preserved. A commit can be signed based on its SHA-1 representation, its SHA-256 representation, both, or neither. The bijection means the two types of signatures are in a sense interchangeable, or in other words the bijection between object representations implies an equivalence relation on signatures. An SHA-256 converted repo can quickly validate an SHA-1 based gpg signature using the lookup table.
If we are changing to sha256 because sha-1 is broken, how can we trust all those mappings and signatures?? This makes it even more confusing as to why this change is being made
Git moved to a modified SHA-1 (let's call it SHA-1′) back in 2017 which prevents the particular attack described in SHAttered, so the SHA-1′ object names and signatures can be trusted for now. But the known collision attack on the underlying algorithm makes it more likely for SHA-1′ to be broken in the future too. Researchers could find variations or improvements of the weaknesses used in SHAttered to attack SHA-1′ as well. In other words, there are no known practical collision attacks, but there are known avenues of reasearch that could lead to practical collision or second-preimage attacks being discovered.
The idea, then, is to migrate all git history to a stronger hash before that happens. If we waited until an attack was found, it'd be too late; all git history would be suspect forever into the future (disregarding extensive auditing), even if a stronger hash was used retroactively. Full SHA-256 doesn't have known practical collision attacks, and it doesn't even have known research avenues likely to lead to practical collision attacks, so is considered more future-proof than SHA-1′.
This idea of future-proofing is common in cryptography. For example RSA-1024 keys are now considered practical to factorize (with a supercomputing cluster and a few months), but that's okay, because, foreseeing the possibility, "everyone" switched to RSA-2048 or stronger a decade ago. Now we're adopting new key algorithms to protect against possible quantum computing attacks in the future.
Yes, I understand all that. The original author points out that completely ditching sha-1 will be a huge pain, very costly, and won't meaningfully improve security. You responded by saying we aren't completely ditching sha-1, which seems to lend credence to the argument that sha-1 is still safe enough to use. That's where I'm confused. Having two hashes for each object in a git repo, a sha-1 and a sha256, is not the same as wholesale deleting your rsa-1024 keys and replacing them with rsa-2048. If I'm only looking at the sha-1 and all my tools only look at that, then the sha256 isn't helping me. Is it?
Ah, the plan is to eventually completely ditch SHA-1, but only after "everyone" has already converted to SHA-256, which is hopefully before SHA-1′ is broken. See sections Object names on the command line and Transition plan from the link in my top-level comment.
If you trust the sha-256, you know the commits you have and their ancestors are not evil. So, the sha-1s for those commits can be trusted.
You should also check the mapping table's sha-256 hash, to know that it's not evil.
If you later get an evil commit trying to masquerade as one of those, Git will presumably notice that there are two commits with the same sha1 in the same repo, and bail loudly.
I threw in an upvote, there is previous history in software for version migration. With version 2 to 3 being the migration here, surely there's a way to stuff 2 bits of info somewhere in a git dir and just patch the version 2 today to start looking for those bits.
now I'm confused: if sha-256 is only used for new repos (default) and sha-1 continues to work for sha-1 based remotes, so a git clone, git push just keeps sha-1, then what exactly is the remaining issue? to me this sounds like a perfect and user friendly migration strategy.
The third-party tooling concerns seem plausible. For example, it seems like the intension for Git hosting is:
1. Git users all transition to SHA-256 while git hosts stay on SHA-1.
2. Once almost all Git users are on SHA-256, Git hosts flip a switch and convert all hosted repos to SHA-256. Git users on SHA-256 don't notice anything, because the only thing that changes is how the client and server negotiate which objects to send.
3. Git hosts disable SHA-1 support. Git users on ancient clients need to upgrade or switch hosts.
Instead, by the author's account, Git hosts are deciding to convert repos to SHA-256 individually and placing the decision of when to do each conversion upon their users, who may not be informed about the situation. To me, that seems like the wrong decision, but then I don't run a Git hosting service.
No doubt other tools will take some time to catch up too. For example, libgit2 is perhaps the most popular third-party Git client, and it still doesn't support newer Git features like reftables (circa 2020). Though libgit2 does appear to implement SHA-256.
One of my favorite fun facts about Fossil SCM (another source control by the devs of sqlite) is that they patched their use of SHA1 6 days after the shattered attack was published:
"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."
To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.
That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
It's easier to make world breaking changes when the world is really small. If git could magically just get everything and everyone to cut over and use git 3.0 in a magic instant, it wouldn't be having this problem.
It seems Fossil can use multiple hash algorithm within the same repository. Yes, that decreases overall security since the hashes pointing to older objects can still be spoofed.
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.
Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
Also, interestingly, Git today does _not_ use a straight SHA1 because of these attacks. It uses `sha1dc`, a slower collision detecting variant that specifically checks for this vector of attacks. So currently, Git's SHA-1 variant is not susceptible to the SHAttered/Shambles attacks.
Sure but that's just kicking the can down the street. It will eventually be susceptible to a collision attack. As will SHA256 btw. What then? Are we doing to go through all this again?
Online video has handled this. There are various codex, container formats and transport protocols. The TLS handshake does this. The ability to deprecate and replacing the hashing algorithm should've been built in from day 1.
As an expansion on that, as far as I know there's only one implementation of Fossil. There was at least a Java one for git, I believe, and I don't think upstream git is related to libgit2 either. And GitHub is doing… something not visible from the outside.
(… looking at the parent, though, I imagine there might be some information from the inside…)
> As an expansion on that, as far as I know there's only one implementation of Fossil.
That's is, since only recently, no longer strictly true: the age-old libfossil recently got client sync support, so is now (aside from _serving_ repos) essentially a standalone impl (its own developer still uses fossil(1) stash, patch, and diff -tk features, but otherwise uses libfossil's counterparts).
Also, Dan Mestas has <https://github.com/danmestas/go-libfossil>, a Go library for working and fossil, and he is also working on <https://zeitforge.app/>, a clean-room impl. of fossil (whereas libfossil is largely ported directly from fossil(1)) which even goes so far as to _not_ use an sqlite database for its file storage.
Dan Mestas and Dan Shearer are working on finalizing RFCs for fossil's sync protocol and artifact format, and zeitforge is created by carefully managing LLMs which are reading that draft (but not the source code of libfossil or fossil).
I personally have been using `go-libfossil` in a project! It's got some sharp edges but insanely useful in avoiding CGO and using fossil at the same time
To me it's more of a reality of creating a tool with a huge active community and a community of contributors and creating a tool with a small team and small community.
Isn't sha1-dc just checking for a small list of hard coded disturbance vectors? Meaning that script kiddies can't easily reuse exact values that researchers spent a bunch of time precomputing. That's far from being a good long term solution.
I'm not an expert on this but my impression is its a bit in the middle. It isn't a great long term solution, but finding new disturbance vectors is very hard. Its not like someone can just go find more on a whim.
No it’s different. Git supports both but you cannot mix them in a repo. Fossil allows mixing them in a single repo. This is clearly stated in the link.
I always liked the ipfs concept of a multihash where a hash algorithm id is stored with the hash, I don't know how well it worked in practice, but in theory all applications of the protocol are now forced into a world where there are multiple hash formats and it can and will change.
We don't need that. Introducing an unknown hash algorithm itself is a security issue.
This problem isn't hard or new. Just look at things like a TLS handshake. You need to separate the protocol from the storage implementation.
I personally believe the initial Git programmers were too in love with the efficiency of doing a bitwise 160 bit comparison on the stack and they sacrificed the known issue of changing the algorithm to do it. A decade earlier the same thing had happened with MD5.
> but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, Okay, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get. ... [1]
It's of course possible that SHA-1 was not originally intended as a security feature, but according to Hyrum's law, every observable API behavior becomes something that somebody starts to depend on, so if you start very publicly shipping a cryptographically secure hash, you better keep it cryptographically secure.
At the very latest, this fact was cemented when first-party git commit signatures started depending on the security properties of SHA-1.
> if you start very publicly shipping a cryptographically secure hash, you better keep it cryptographically secure
So if my API happens to return text strings that always happen to have an even number of characters, I better make sure that all future versions of it also always return an even number of characters, just in case some moron decided to bank their application's functionality on that? No. If you decide to write a fragile application tethered to some incidental property of some upstream software, your application deserves to break.
> No. If you decide to write a fragile application tethered to some incidental property of some upstream software, your application deserves to break.
Just because you never made any promise regarding one aspect of your API doesn't mean that you're absolved from responsibility when you choose to change it. If you know for a fact that many users rely on it and you choose to break it, you need a good reason. That kind of balancing act is part of your job. If you don't respect your users, perhaps development wasn't the right career choice.
As developers we're of course constantly tempted to rename things that we named poorly, or change a schema that is no longer optimal. But we must always take a step back and think about the downstream impact.
> I better make sure that all future versions of it also always return an even number of characters, just in case some moron decided to bank their application's functionality on that?
Oh yes, this does happen. There's even a name for that: ossification (https://en.wikipedia.org/wiki/Protocol_ossification). You can't change your API/protocol, because "some moron" started depending on implementation details.
It's easy to say "your application deserves to break" from an ivory tower, but it's often not easy or viable to fix it (for instance, it might have different owners, it might no longer be maintained, it might be more expensive to change, etc). And the change which broke that application was not in it; the blame naturally goes to what was changed last.
Hey, I'm just the messenger here, if you don't like it, take it up with Hyrum ;)
But seriously: If you can afford to break your user's applications if they "deserve to be broken", sure. Many API maintainers can't, or at least don't want to.
In the latter case (which is honestly the norm rather than the exception, at least for public APIs), yes, you should better think about all implicit API contracts your API shape might be projecting. That guideline has served me very well through my career, at least.
> It's purely a consistency check. The security parts are elsewhere
Sorry if the video answers this, but how does commit signing work if it doesn't rely on the hash algorithm being resistant to at least second-preimage attacks?
It was a mistake to assume a fixed algorithm in the repository format and client-server protocol. I remember being surprised when I learned about that choice, being familiar with cryptographic protocols and formats where the hash algorithm is usually a parameter that can vary for each concrete hash.
> It was a mistake to assume a fixed algorithm in the repository format and client-server protocol.
See also perhaps Wireguard, which touts itself as not having "cryptographic agility" because they wanted to avoid all (perceived) problems and complications of IPsec. But now that PQC is (allegedly) approaching there's no easy to update things because (AIUI) there's no negotiation possible in the protocol; you're basically standing up a 'Wireguard 2.0' that runs separately than the original.
Yeah, in 2005, SHA-1 was just about the best you can do given the constraints of the time (without picking something much more esoteric, much slower, etc.). Using SHA-256 at that time would have been noticeably slower on the computers of the time, and made repo metadata take up a lot more space.
Have you seen algorithm-agnostic protocols, like TLS v1.2 and IPsec? It always turns out to be a bad idea. Future upgradability is fine but you really don't want implementations of the protocol to be incompatible with each other and you don't want attackers to have any chance of tricking you to using insecure protocols.
Flexibility made sense in 1995 when nobody was sure which algorithms would stand the test of time. Even in 2005 it was unnecessary and in 2015 it was an outright liability. If you have a good algorithm just specify the good algorithm, don't let the parties negotiate either a good one or a bad one.
Your argument is weird because it assumes we can know if a hash function will be secure forever. Assuming git would use SHA1 forever seems shortsighted.
That was true in 2007, but it stopped being true once people started signing commits and tags. A GPG/SSH signature on a commit covers the commit object, which names its tree and parents by SHA-1, so the signature is exactly as strong as SHA-1's collision resistance. If someone can prepare two trees with the same hash, a signed release tag vouches for both of them.
SHA-1DC blocks the known SHAttered/Shambles-style attacks, which is a good stopgap, but it detects known techniques; it isn't a hash you can reason about.
Code signing went through the same thing. Authenticode signatures with SHA-1 digests are effectively distrusted on Windows now, and that migration hurt precisely because it was put off until it was urgent.
Linus is one of the last remaining champions of rationality in large-scale software projects. Otherwise, it’s filled with devs who love to leave their brains out when it comes to practical scenarios. Anyone who thinks there is a security issue here is an absolute moron.
Rationality is the key point. A lot of devs approach their work in ways that are not rational or practical. Like the dev who wants to build a gold filigree decorated elevator that can handle ten thousand pounds of cargo and moves with the smoothness of a magnetic levitation rail in order to reach the second floor when all you really need is a ladder for the 2 people that will need access.
I don't understand why Git is not making the SHA-1 and SHA-256 modes far more compatible with each other.
SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.
But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.
And now it would be impossible to get a new SHA1 collision in to the repo.
The only new git features needed would be:
a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object
b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit
c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
They begin by saying mixing can never work but don't address the option of dual hashes... Then the interop plan is basically a secret git format that maintains both sha1 and sha256 hashes FOREVER.
I don't see why all users wouldn't want to keep both hashes.
I skimmed the video, and I didn't quite catch that. Near the end of the video, however, she did mention that interop is in the works[0].
In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.
Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.
My point is not that it's the end of the world (or the end of Git), but that it will be painful and unclear and confusing to lots of people. That would be fine if it made a huge difference in trust or protection, but it's the wrong way to do that.
Hm but the date is stored inside of the commit. The only way we can know that a commit's date is authentic... is through its hash. If I can forge commits with any SHA1 hash at will, I can make a repository whose head commit has the same SHA1 as the one in torvalds:
/linux but where any commit was replaced by a malicious commit with the same SHA1 and a fake date. You have no way to detect that my repo is inauthentic other than through a deep history comparison. The whole idea behind a merkle tree is that just checking the hash of the top is sufficient to know the identity of the whole tree.
I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.
If there is one way enforcement (i.e. there is one point where the last SHA1 commit was signed by first SHA256 commit), I think it should be safe ?
The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered
Every file (indirectly) referred to by a SHA256 commit using a SHA1 hash in some tree object can still be spoofed. Fixing that requires rehashing all objects and recreating all tree objects so the tree objects referred to by SHA256 commits are purely made up of object references computed by SHA256.
The date that a repo receives a commit is known to that repo. And a repo can stop accepting new SHA1 objects. And a SHA256 object could have a flag that says that no SHA1 objects may ever reference it.
The design of Git, as a Merkle tree, is meant to allow for use cases like this:
* I host a mirror of the Linux git repo.
* You download Linux from my mirror.
* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).
* You check whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.
This is a completely normal use of Git. People download from mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.
If I can forge my own commits to have any hash I want, this whole model breaks down. I can replace some old commit in the repo with my own forged commit with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.
Do I understand correctly that GitHub currently doesn't support SHA-256 repos at all, so the author faked a screenshot of what the create repository UI would look like if they choose a really stupid way to add support for SHA-256 repos, and then proclaimed it an unsolveable problem? Selecting the repository format based on the first thing pushed to it is really not a crazy idea.
The submodule problem is real, but the forge problem is entirely just "if forges implement support in a way that makes it painful it will be painful", and that's true of literally every feature that requires forge support.
I did generate that screenshot as I am not in the beta for this - I don't think many outside of GitHub are. However, this is exactly how GitLab does it and I would be _really_ surprised if this is not almost exactly how it's implemented. There is almost no other reasonable way to do it.
You can't select the repository format based on the first thing pushed to it, because both sides have to be initiated before a transfer can happen.
So you either start the project on the server and then clone an almost empty repository to start working (which I think is rare) - in which case you need to choose the format like this.
Or you initialize it locally and start your project and then push it to GitHub at some point, in which case you need to initialize a server side version that matches the format to push it to. There is no "initialize a new thing on push" and there never has been.
I personally would have preferred if there is a proper disclaimer that the screenshot is an approximation of how the UI would be. Or alternatively, use a more distinct art style of conveying the UI (e.g. comical/sketch lines or whatnot).
Right now its only a vague statement of "will need to look something like this", which does not imply on the originality of the image. I myself am misled that the screenshot is a legit UI.
> both sides have to be initiated before a transfer can happen
What would be preventing a forge from answering with a reference to schroedingers octocat when inquired about the yet-undefined properties of a newly created repository? The first pushing client explicitly looks for it, and no final decision has be made server-side until someone wants to take a peek, no?
I also don’t know how big this problem really is. You practically have two ways of creating a repo. Local first and push or create a repo remote with Readme etc and clone.
The submodule issue is a different beast. But that one is currently a problem as well when a person choose a http url or ssh url. I have some extra git configs to normalize everything to ssh for instance.
but yes: that ui in the article is likely entirely imaginary, but it's likely an option (either in the ui for private beta users or by raising a support ticket or by some internal tool) to get sha256 repos today.
I completely agree with everything you say here. I felt the critique in the article was way more strawman falacy than everything else. I use(d) git much more locally, and would never run into this problem because I would only use sha-256 on new repos. Rewriting the entire history is also not so much of a pain depending on organisation size and how it is done (if an org decided to do it). It may be a pain for git-hosts, but even there i think reasonable solutions can be found. Still not sure if I care about the switch though sha-1 worked and the security implications seem minor to my use-cases. I am much more concerned with code leaking online...
sidenote: i hate the term "forge", is it too late to settle for something less... cringe??
I think this change is more to do with politics rather than "security". Those kind of things where companies or gov, need to be certified with those super secure certificates and can't be using software that uses SHA-1. I don't have proof, but I'm not doubting it either.
This is what I saw in one of the mails.
> > There are organizations where SHA-1 is blanket banned across the board - regardless of its use
And also on git 3.0 breaking changes.
> > SHA-1 ... recommended against in FIPS 140-2 and similar certifications
Since SHA-1 isn't used for security in git, they should've instead moved to a non-cryptographic hash function such as MurmurHash3 and avoid all these problems, instead of moving to SHA-256 until SHA-256 is broken and need to move to the next cryptographic hash that is now incompatible with previous versions of git repositories.
SHA-1 is used for security in git. It's the thing that guarantees a commit SHA is unique. Without that, you open up all sorts of downstream infrastructure to supply chain attacks, where old objects get replaced with malicious ones, and then replicated on each subsequent git pull.
Linus' old argument was that the substitution would probably be noticed eventually, but that's specific to the way Linux uses git, and what he said probably isn't true in practice -- even if it is, there have been enough supply chain attacks since then to prove that even temporarily serving the wrong stuff to developers or CI is enough to allow lateral movement into other packages, production machines, etc, etc..
> At that time, Torvalds responded that SHA-1 is not the real security mechanism used in Git and, as a result, even a full compromise of the hash function would not necessarily be a problem.
It seems that the usage of SHA-1 is interpreted as a security mechanism while Linus used it mainly for other reasons, such as look up speed and deduplication of objects.
You can read the original README file when Linus created git[1]:
>+TRUST: The notion of "trust" is really outside the scope of "git", but
>+it's worth noting a few things. First off, since everything is hashed
>+with SHA1, you _can_ trust that an object is intact and has not been
>+messed with by external sources. So the name of an object uniquely
>+identifies a known state - just not a state that you may want to trust.
> ...
> +Another way of saying the same thing: "git" itself only handles content
> +integrity, the trust has to come from outside.
Yes if SHA-1 is broken, then content integrity can be broken but to me it looks like Linus at the time looked it from the point of view of corruption of files instead of "malicious" files.
> There are organizations where SHA-1 is blanket banned across the board
This is very likely the case. And if it is, then it's a lost battle. You simply can't reason with that kind of corporate people, let alone have an argument around this level of complexity. Kafka (the writer, not the message broker) predicted this 100 years ago.
When going through the article, my instinct was changing from "annoying" to "this really sounds like a Python 2/3 moment for Git" to finally "oof this is going to be a mess" in the libraries/submodules part.
In my experience corporate box checkers don't care about reasoning (much less "encryption") at all, they see it as an annoying blocker in their path to the next promotion.
The difference with IPv6 adoption is that the internet relies heavily on network effects: so long as some hosts only have an IPv4 address, you need an IPv4 address for full connectivity, but then if everyone has an IPv4 address anyway, there is no immediate need to migrate to IPv6.
(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)
This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.
If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
Github is THE main platform for git.
If github doesn't upgrade (and their code has been shit and hard to fix/update before) then the shift will not happen. Because yes, you can upgrade your repo independently, but if there is nowhere to push, no one will do it.
IPv6 is (also because of github) a great example for this. You can easily have an IPv6 address next to your IPv4 address, but a lot of websites (e.g. github) don't have that. Why would a normal company use IPv6 if even the bastion of nerds doesn't use it?
And to Python 2's "eventual migration" I can unhappily tell you, that my company (recently) bought an actively developed tool, that still uses Python 2.
I don't think GitHub is quite as important as that, outside of some specific projects that made bad/lazy decisions (thinking Golang here). If they didn't support 256 I think a lot of enterprises and open source stuff would simply jump ship to GitLab or elsewhere. In the enterprise world it only takes one person to write a security document banning sha1 for this to happen. For that reason, GitHub will support sha256.
FWIW there's nothing about the addresses themselves that stops you subdividing a /64, but it depends on your router.
You could put your server at x::1 and static-route that address as a /128 on your router, if it supports it. The reverse route might be a bit tricky but putting ::0 on the router and telling the server it's a /127 should work. Anything outside of the /127 (so, all the randomly generated addresses on your home network) would go back through the router.
Now if that /64 is also changing every day, then it's a problem and idk what you'd do.
I just started learning IPv6 with AWS since they charge $0.005/hr per IPv4. Maybe it will be more expensive in the future and eventually it will be the new default.
It could also be another Python 3 situation: backwards incompatible, unclear benefits with many downsides (3.0 and 3.1 being very slow), large number of libraries that need to catch up.
Does anybody know why it wasn't implemented in a backwards compatible manner?
One could wrap a whole merkle-tree with an additional extension tree, that just adds the new hashes. That way both kinds hashes could be used to traverse all data. The new hashes could be used to check the consistency, the old hashes would still be there to use in UIs or old release documentation. The downside being that you introduce more nodes in the overall data-structure which will have to be supported basically forever. And if SHA-256 is to week a third layer would need to be introduced. But the point is, it could be done. Albeit it would loose some of the elegance of the data structures involved.
Allowing both hash algorithms is, from a security standpoint, equivalent to just using the less-secure hash algorithm.
All repos need to end up using SHA-2 exclusively by the end. All tools that speak only SHA-1 need to be made incompatible intentionally. If the SHA-1/SHA-2 hybrid approach could allow that to happen, then it would be useful. If not, then it would just be a waste of time.
The way you would do that generally is to make it backwards compatible, then adding warning to legacy tools, then turning SHA-1 off by default, then removing it entirely.
Doing it in a backwards incompatible way creates a chicken and egg problem, can't convert repo to sha-256 because some tool doesn't support it, tools don't have an incentive to be updated because no repositories.
Is it? There are many benefits in principle to IPv6, but if my ISP continues to assign me a single dynamic IP, those benefits are entirely moot for me.
Right, but just because you don't happen to have IPv6 right now, how does that remove the benefits for others to have IPv6? That's like saying having a faster CPU wouldn't mean faster performance, because I don't have that CPU yet.
There's a massive push right now from top down to have secure software supply chains. Google SBOM and SigStore. It's not an organic need but if you have government customers you don't have many options.
I thought this would be a snark but it's an extremely well put together argument against the "Hashmageddon".
If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.
I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
Technically, git's design (thanks to very smart people trying to solve this problem like brian and others) is _very_ easy to modify to different hashing algorithms now. A lot of amazing work has gone into this in recent years.
However, it's not a git problem. It's an ecosystem problem. It's that every git repo has to choose one and they're entirely incompatible with each other. That is the cost and the difficulty.
I think there is a question though when that will happen and if it will be in our lifetime. SHA-1 started showing weakness in 2005 (collision in 2^69 instead of expected 2^80. This was later brought down to 2^61 in 2011), the same year git was invented. Nobody has found a similar weakness in SHA-256 as of yet. SHA-256 is still at its design strength of 2^128
It took 20 years to go from vulnerability in sha-1 to having to replace it out of caution. There is no such vuln in sha-256 yet. It could easily be 25 years before we find one, and another 25 years before we have to do something about it. Perhaps longer. Will git still be used 50 years from now?
> All it takes is just one collision to consider it broken right?
No, its considered broken before that stage. i.e. when someone discovers an attack that would allow someone to create a collision faster than they should while still being impractical.
> With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.
Computer power doesn't super matter, what matters is algorithmic breakthroughs. So far i dont think there are any examples of major breakthroughs of that type via AI, although perhaps i am just misinformed. Its still early in the AI revolution, it might still happen, but as it stands i don't think there is any reason to worry about that.
It does though, even if it shouldn't be blindly taken as gospel. Arguments help, but some folks have a better brand of apple box to stand on, and "I live and breathe git" helps quite a bit ;)
Tools like git-filter-repo[1] support rewriting commit hashes in commit messages. git-filter-repo actually does it by default; see `--preserve-commit-hashes` in the manual[2].
... do not migrate old repos? I'm not sure why people would do that. Or, if they do, why would they replace the current repo name instead of creating a different one and keeping the old one closed to make the references work.
I don't think this is going to be a problem at all.
Note that there exists multiple ways to continue to lookup SHA1s in a SHA256 repo.
One example is to maintain git-replace refs for the rewritten SHAs but there also exist config flags to enable object format compatibility extensions that help translate the SHAs back and forth.
Once this starts being actual pain, we will each vibe the replacement index creator (git already supports replacement objects), for back-forth conversion, populated on pack and object indexing.
For massive perf and mem use damage. But oh well. And then we will wait for official version
There are plans to keep sha1s around in a database, but as far as I know, no way to transmit those, so they seem specific to individual forges. They can be recomputed, sure, but again, any signatures break and it's possible that in the case of an actual replacement, the recomputation is now wrong and not easily comparable. So what is the point?
This is not about forges, it is about the repo I have in a folder on my computer. The sha1 hashes shouldn't go away. Yes the forges also need to support this.
Tbh, I would have imagined the SHA-256 transition to work differently.
Let's treat git as a SHA1-keyed object storage. The problem is that we currently use SHA1 both as database key and as hash for integrity validation. At first, I would have only changed the latter.
Local: Request object with key x (SHA1). Remote: Here are the bytes for key x (SHA1) with hash SHA-256. Local: Validate the bytes vs SHA-256.
Local: Store the following bytes with key x (SHA1). Remote: Check if key x has ever been stored in the database. If so check that the SHA-256 of the new bytes matches the SHA-256 stored under key x. The only thing the repo has to keep is a LUT from stored keys (SHA1) to hash (SHA-256). The prevents SHA1 collisions from being stored.
The SHA1 key then just becomes a convenient alias for an object. The only restriction is that you can't have two objects with the same SHA1 in a repo. We already kinda do this when we refer to commits with the first few chars of the SHA1. If there is a "collision" git already detects it and asks you for more characters.
Finally, you can convert the internal representation of the tree to SHA-256. If some legacy client requests aliases via SHA1 you use the LUT, new clients request the SHA256 directly.
I positively don't care about the collision issue.
If you need to certify the authenticity of some code, and you've decided that a Git hash of any kind is going to be your certificate, you have a problem between keyboard and chair which is not fixable by stronger hashes in Git.
It was about 18 months back, with gitea, I was starting some new projects and went sha256 because I’m a nerd that adopts things early. I could do basic git stuff, but I was stunned by how much didn’t work. A lot of actions, maybe even the entire action runner itself wouldn’t/couldn’t/didn’t work. It’s a deep cutting change and an expensive one that doesn’t have a new feature value. I abandoned that effort and rebuilt my repos.
Intentional attacks and collisions aside, what do you say to the fact that Sha1 had a lifespan at design time? NIST has issued retreating usage guidelines for the last 15 years and they themselves say it should be completely phased out by 2030. It just seems like good hygiene to switch it out. I agree, it will have a long tail and be a big bowl of suck, especially for the not dead but rarely touched code.
I appreciate the intent of “modern” tools with no switches, no ways for devs to make bad security choices because the hash, cipher and encryption mode are fixed vs all the crazy looking asn.1 stuff in tls and openpgp to enable n-degrees of configuration but this very issue is the counter example.
So they’ve been talking about this for many years, planning, and finally announce when they’re going to switch the default.
So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?
I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.
This seems like a bunch of Monday morning quarterbacking.
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
Your argument is persuasive and well illustrated. I think the problem is the intro paragraphs come off as too certain of catastrophe which, when juxtaposed with your claim that "smarter people than me have been working on this", makes it sound like you don't actually believe they're smarter than you. The rest of your essay feels fair and not judgmental.
I do believe they're smarter than me, but sometimes very smart groups talk themselves into ultimately impractical solutions because they're all smart. Sometimes you need a dumb guy to come in and say "are you sure this is right?"
It's a hitchhiker guide to the galaxy reference, where sure, something is technically available but not clearly published and there are hoops even for those who know what they're looking for.
(No clue if it's applicable here, I'm not aware of this case, but I believe that's the reference if it helps :)
Edit : exact quote, as Arthur's house is about to be demolished for a highway bypass:
"But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.
I believe it's a reference to the hitchhikers guide to the galaxy – where the plans to remove the protagonists building to build a bypass road was hidden in this way.
nofunsir is a stoichastic parrot, matching to a bit in Hitchhikers Guide to the Galaxy, wherin the protagonist should have known to protest a plan to demolish his home where plans where clearly documented in a hard to find place that they could not have known about. It is not a good pattern match, because git has been discussing this in public on documented mailing lists for years.
A number of years ago when I heard about this, I was pretty angry and made a private fork of git immediately in which I tried to scrub away the SHA-256 bullshit. But that's basically just paddling upstream with a spoon for a oar.
The stewards of Git are going to do whatever they want, and there is nothing you can do about it if you don't have the clout to create a fork that takes the lead.
No amount of discussion will do anything because they've already decided that their view of the situation is correct. Git hashes are not just content identification but a digital certificate mechanism, and their collision resistance is a grave issue that must be fixed, the end.
You will be browbeaten in any discussion; it's not worth the energy in a world replete with issues.
This is actually a good change. If you want to change the security assumptions of Github repository, ie make them somewhat distributed. Then the SHA-1 based commit hash is a major problem. It only costs about 10k in 2024 to find a collision to a random SHA-1 hash. While this costs essentially makes the attack infeasible for most threat models. It does limit how far you can scale this without an obvious footgun waiting for you.
This is a good change, even though there is a massive technical debt in changing such a widespread system. It is worth the effort. Should generations from now still be using SHA-1 for their Git ops? Sometime you have to do the switch, otherwise you will never progress.
For my usecase basically i needed to know that every Git commit pointed at a cannoical blob. With SHA-1 you could generate two blobs which hash to the same SHA-1 hash, while you can do the format verification which helps i could not do that in my usecase as i did not know the underlying data. To fix this i had to very ugly have two methods of referencing any Git object, a cryptographically secure SHA-256 ID and the Git ID SHA-1.
Yes every repo is either one or the other but you fix that by rehashing the entire repo. Everyone can do this independently. It's entirely possible to maintain to identical repos in SHA1 and SHA256 mode but for the most part I suspect once updated people will simply pull down the new repo and use git 3.0 as a required version.
As migrations go, it's reading as simple to me. You'll just have to backpoint the commit signatures. I must assume there's a backwards compatible reference for them in git 3, right?
Or drop them and reference the old structure in a dire pinch.
Re-hash the entire repo as in rewriting all history? Hooo boy will that be a mess, I deal with things which reverence commits by hash in repos all the damn time. There are thousands of them in every Yocto project!
Do you have any external references to any commits that matter, for example in your communication platforms (emails, Slack) or your bug tracker? Or, worse yet, in places where they aren't just text format references, but used for things like CI/CD caching decisions or security scans?
Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
I think I remember something for this for mercurial to git migrations, I hope when its git 2 to 3 something similar is made (or it will take some time to adapt like when python did its 2 to 3 migration).
"The migration to the new format is simple; Just re-write everything in the new format, but also keep the old format around forever too since data is lost in the new format!"
>it will be an incomprehensibly expensive and ultimately valueless and avoidable global nightmare.
thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
Actually, I think sha-256 is possibly faster than the sha1dc variant that Git currently uses.
I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.
How could SHA-512 be faster? It does more rounds of the exact same operations as SHA-256 with a bigger state. Although if it really were faster, SHA-512/256 gives you a truncated version.
The problem is existing repositories have SHA-1 commit hashes, and if you also end up changing those...
There's lots of tools that refer to git commit IDs. Some of those tools may even hardcode a commit ID to be 40 hex digits long. The fact that these tools are external also means that "oh, just rewrite the commit messages or code to refer to the new IDs" isn't feasible. The only way to not break the world is to let people refer to existing commits with their SHA-1 hashes in perpetuity, and it doesn't sound like git is set up to allow this in any way, which means that existing repositories have to stay SHA-1 in perpetuity and that will cause fun down the line if you start having to make SHA-1 and SHA-256 repositories.
Changing from master to main is a one-off change. It might require changing your scripts once to refer to 'origin/main' instead of 'origin/master', but other than that, there is essentially nothing more that needs to be done, there is no risk to historical artifacts that needs to be mitigated.
> Mathematically, for SHA-1’s 160-bit output, the birthday bound means that you would need about 1.4 septillion random files (1.4 quadrillion billion files - 1,400,000,000,000,000 billion - it's impossible to effectively describe) in a single project to have file hashes accidentally collide.
The birthday problem gives the number of hashes to have more than 50% probability of having a collision. To have a collision, only 2 hashes are needed, with very low probability.
I don't agree that SHA-1 is much longer feasible for git.
But I also don't think that switching to SHA-256 must be painful. A git2->git3 converted repo could just store all the past hashes, so existing links don't break.
I was thinking the same. Can't the git CLI see a hash and say "well, I don't see any matching SHA-256 hash, but let me check the Legacy SHA1 hashes I have stored", and still resolve an old SHA1 hash to the correct commit?
It's Python 2/3 or Node CommonJS/ESM all over again. I thought the lesson of history was that it's more reasonable to live with a known problem (and have mitigations for it) than to "fix" it in a way that breaks everything that's good.
Prior to SHA1 we had MD5, a decade earlier. MD5 collision attacks had already been widely documented and known. It was the most obvious thing on Earth that this would happen to SHA1 too. Apparently, Linus never realized there was a need for cryptographic security and that the hash was purely internal.
Here's what I honestly think was a factor. I think C programmers fell in love with the implementation that you could throw around a fixed hash record on the stack. It's incredibly efficient. But it's an efficiency that doesn't really matter because as soon as you read from or write to a disk or a network or even memory, any cost saving is completely gone.
More than a decade ago, some people wrote a Java implementation of git (jgit?) and despite all their optimizations, it was (IIRC) only half as fast as C git. It is of course because Java at the time had no concept of stack values for non-primitive types so couldn't compete. Personally, I was impressed: only half the speed? That's pretty good.
For something that's only 20 years old, the Git SHA1 assumption is some of the worst technical debt we have in the modern era.
Here's another thought: when people make a lot of these programs, they often make the mistake of not separating the program version and the network protocol (or just the external API). So you end up with brittle client-server implementations where you have to upgrade both the client and the server at the same time because they lack a network abstraction.
The other end of the spectrum is video streaming where you have codex, container formats, transport protocols and so on.
Having gone through multiple "security" upgrades over my decades-long tenure, I firmly agree with schacon here (this post is mostly for schacon, since I'm reading a lot of negative on this discussion (Rubyists need to stick together)).
The hashing algorithm's security properties are a security property of Git. This is because:
* we pin to commit hashes and expect this to refer to immutable content,
* commit and tag signatures are over the hash.
The Linux quote about trusting the distribution doesn't make sense to me, as Git is content-addressable and decentralised, although possibly at the time it was a reasonable position to take for kernel development, but [1] could equivalently happen for Git and the hash algorithm being non-broken is required for it to be noticed. Not having to trust the forge is a very desirable property.
I think this is really selling the severity of "SHA-1 is a Shambles" short. The author distinguishes between second preimage and collisions, but Shambles is a chosen-prefix collision, so it's sort of in between the two ends of the spectrum he explains.
The difference here is between me being able to send you a benign file and then swap out a safe file (the scenario the author describes) and me being able to swap out your own familiar, benign file except this time it has malicious content at the end (along with a bunch of garbage).
That second scenario is obviously much more dangerous from the perspective of human review and noticing that something is wrong. If you have a large, rarely changing file already, then content silently smuggled onto the end of it could stay undetected for a long time.
(In practice, it's not that simple; you'd have to actually figure out how to make multiple pieces line up with chosen prefix collisions, and I haven't worked through how you would do it. Maybe it's impossible without further weaknesses. But that guarantee is feeling pretty threadbare.)
Anyway, this doesn't matter because git switched to sha-1dc back in 2017, and afaik, this addresses the shattered/shambles problems. I find it odd that this isn't the point the author focuses on; instead, the only mention of sha-1dc is in a footnote saying that it could be replaced with his proposal.
And sure, maybe that's true, but why isn't that the entire argument of the article then?
Well, this could change at any time. Part of the point was to argue as if SHA1 was completely broken and I still feel that its the correct argument.
But from your chosen-prefix argument, having some object thats been around a while is a problem, because any viable attack needs to be fairly new, since git wont replace objects it already thinks it has. Maybe fresh-shallow-clone scenarios like GitHub actions, but that still always has a non-sha based authentication protection (in other words, actions never run on untrusted code and that trust is never based on signed artifacts but on source provenance)
akschually! I argue in a footnote at the end that sha1dc is an unnecessarily expensive shim protecting codebases in a similarly unnecessary way and should be removed so that our pushes and fetches can be much faster.
With the increased computation given to LLMs for the new development practices, wouldn’t any extra sha be a much smaller accommodation (then claude, chatgpt, etc…)?
The post's argument that hash collisions are irrelevant in practice is not convincing at all. Basically they amount to:
1. Collisions aren't as bad as preimage attacks
2. Even if you made a file-with-malicious-hash, how would you get people to pull it?
3. Other attacks are a bigger problem (social engineering)
(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.
(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.
The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.
(2 counter) is impractical because all nodes of git will not replace objects if it thinks it already has it. So any attack has to assume this is the first time the node fetched, which is difficult before trust is established, which is difficult. This is part of the argument Linus originally outlined for this vector, which is that it only works for _very recent_ objects.
(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.
The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.
It doesn't reject, but it will not replace. Same for a fetch/pull. That is another issue with this attack vector (that Linus also mentions) - it has to be the _first_ time that a node has seen this object. It makes the attack even more difficult than it already is (in like 4 different major ways)
In the middle of vacuous complaints about slightly late tools and debatable threat models, I see one legitimate-looking concern in this article: submodules require either hashing algorithm and therefore force inconvenient upgrades.
Is it true? Where exactly the parent repository references hashes from the submodule repositories, and how could these links be generalized for compatibility?
Ugh, I didn't know that SHA-1 submodules wouldn't be supported in SHA-256 repos. That changes the transition from painless to a major dumpster fire. Having to maintain converted forks, and use different hashes from upstream is going to be a mess.
The only reason C++ projects depend upon submodules is because C++ is just as archaic as COBOL these days.
I used to be one of the more vocal defenders of C++ but honestly, if the reason you’re defending git submodules is because there’s no better option in C++; then you’ve basically already lost the argument.
Nearly every other language has found a better alternative. Even Go has, and that’s largely mocked for its prehistoric approach to modern programming language design.
This article says if you manually enable the experimental new code today then “you can't push code to GitHub” and that old and new repos “cannot be mixed”.
The git transition document talks about keeping a bidirectional mapping between old and new hashes, and converting between the two when pushing to legacy repos:
I suspect that’s what gitbutler.com means when they say that, although broken today, everything “will almost certainly be fixed” when git 3.0 ships.
While I’m sure they make a good argument in the hash theory section, I’m less inclined to believe the “train wreck” part of their post.
If you upgrade your house and leave the roof off then yes, it will be a “costly mistake” the next time it rains, but if your plans say you intend to put the roof back on then I’m not sure how I am supposed to interpret a blog post warning of the perils of a roofless house.
Oh, agreed this sounds like a terrible migration path and shouldn't really be needed in the first place.
What I'm missing in the article is whether any Git server accepts replacing a SHA-1 identified object it already has. If it doesn't, then the distribution trust discussed holds, and keeping SHA-1 seems fine. Adding additional signatures seems fine for those who need transitive trust.
> The first thing that you'll notice (other than the much longer hash value) is that you can't push this code to GitHub, though that will almost certainly be fixed by the time Git 3.0 is released. In fact, that's probably the main thing currently delaying 3.0 entirely.
> But when you do want to push it to GitHub (or any host), you will need to tell them when creating the repository on the server that this is a sha256 project. Every project will now be in one bucket or the other and they cannot be mixed.
> This is immediately going to frustrate people because now they need to know what version of git they ran git init with and make sure when they go to GitHub to create the server repository, they choose the right one.
Github can fix this problem, and we have had siilar stuff in the past, when we had to switch repos. This is simple and doable, just need gradual work. I don't see the problem.
I find it disappointing how few people are addressing the proposed "Independent Tree Hash Headers" solution. It's probably the most interesting part of the article, but it's getting the least attention.
I came in expecting to disagree strongly with the article, but ended up agreeing more than I didn't (though I still don't 100% agree, as collisions are still an issue for mirrors). I find the concept of multiple hashes per commit quite interesting. It would allow mixing hashes in one repo, wouldn't break submodules, and tooling could be used to reject commits without any secure hashes for a gradual transition (like enforcing signed commits/tags).
I agree, but you're not going to sell that argument to a herd which has decided that git hashes are digital certificates which must be replaced with SHA-256, or the sky will fall.
Can someone more cyber-pilled than me explain what the actual risk with Git hashes being susceptible to collision attacks is? Obviously accidental collisions are problematic, but to my understanding the probability of that is still approximately zero.
Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a far from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
> who's out here using git hashes as a security tool
Among others, dependency management in Rust [1] and Python [2] sometimes uses references that work similar to https://github.com/rust-lang/rust/commit/ec999ed [3] to suggest one particular version of the project, authored by the specified maintainer.
Unfortunately, it means neither, unless you pushed it. The hash points to whatever the first person uploading it to github submitted. And the author/org name in the URL is window dressing: all the objects go in one big bucket regardless of push permission to one particular fork (because why wouldn't they - today, collisions are believed to be recognizable because the cheapest way to craft them results in clear tells).
[3]: N.B. the "This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository." warning Github has started to add to URLs like that.
Why are we quoting a technical opinion made in 2005 as if 21 years of time passing has somehow solidified the opinion? I mean, maybe it has, but that's not how technical opinions work
I agree with the main thrust of the article, that we should instead trust the transport mechanism rather than the cryptographic properties of SHA1, but because of this I don’t really think switching the default will be a costly mistake. I rarely use full SHA1 hashes right now; I only use the truncated version and I don’t think any user cares about the length of the full hash. As for compatibility with forges, it’s just a small UX problem that should be solvable: don’t let the user choose the format when a repo is created; instead choose it when the first push happens.
I don't why a sha256 project could not support sha1 usages too: if you try to merge/rebase/cherrypick/whatever a sha1 commit in a sha256 project git could simply recompute the needed s256 data and keep old sha1 ids as validated metadata. you could even use either hashes as tree-like refs.
you could even do the same in reverse and "import" sha256 commits in sha1 projects. the only real difference would be which one is used as primary key in git's internal store.
Because it's slower. You would have to build an internal storage format and a network format that allows multiple hashes.
One of my theories (see my other comment) is that Linus and the people who took over Git fell in love with the implementation. A fixed size (160 bit) hash on the stack is incredibly efficient. A variable length record with optional fields isn't. You can do a byte-wise comparison of 20 bytes on a SHA1 hash. You can't if there are unknown hashes in there that might not break equality.
You should never let these kinds of implementation details drive the design at the expense of future compatibility. We were already aware of this issue in 2006. It happened with MD5 (to SHA1). So now we have some of the worst technical debt to be introduced in the 21st century and it's going to be a giant pain to fix. SHA256 will eventually fall by the wayside too so just changing the algorithm isn't the long term solution.
It was theoretical - the point was that maybe some paper is published or some new tech or issue comes up. Now we have to do this again. If we separate the concerns, then we don't have to deal with both as though they're one problem. We can deal with one thing for content addressing and another for trust and security.
Heard. If we’re going through this pain of a migration- adopting something that is forward compatible at the ecosystem/community level like ipfs makes sense to me. You throw a byte at the front of the hash which indicates the format the hash is in.
We have a content addressable storage system internally and we use a self describing hash for it, so we can change the hashing scheme for different use cases.
A lot of replies here seem to be asserting that this "isn't that hard" without addressing the thing that makes it most hard: submodule compatibility and the breadth of tooling
They're the best way we have to reference other repositories from one repository. All other solutions don't have the benefit of being built in to git and having support built in to all git forges.
Ecosystems like Yocto are built around having meta layers as submodules. And, despite the usability flaws of submodules, it works really well.
I also use submodules to include dependencies into C++ projects a lot. It works fine.
git subtree and git subrepo are compatible with all git forges and don't require normal developers to install the extensions. Only the person/bot doing the occasional sync to the external repo has to install the extension. I prefer git subrepo for most (but not all) use cases.
I won't defend submodules, but I also don't accept this as a response because it's irrelevant. They are used and it will be an unbearable pain when they break.
Someone started this FUD a long time ago and it has worked. Instead of using an elegant mechanism, project have built inelegant wrappers on top of git like go.mod which are actual mistakes.
It is the same people who say Git is unseasonably hard to learn. It isn't. Git and its command line is ugly and hard to remember. Understanding what those commands do at a commit and a branch level isn't really difficult. Even the internal object model is only moderately difficult.
I say this as a person who strongly dislikes many many aspects of Unix and Linux due to bad design and terrible UX. Git has a better design than any Unix program you get.
Submodules work okay. It is just Git LFS but for Git repos. Get over it.
In fairness, there's a second issue: Git repos should just support multiple hashes, so you could just add SHA-256 to a SHA-1 repo, which would then transparently support both SHAs.
This even would make the SHA-1 git objects collision resistant when stored on a trusted server, even with untrusted clients. (Exercise left to the reader.)
TFA argues exactly this, that SHA-256 should not have replaced SHA-1, but it should have been added as a second hash.
Nonetheless, this solution has its own disadvantages, which were considered as more important by those who have chosen the replacement solution.
In my opinion, SHA-256 should be used for any new commits, and all old commits should be rehashed, but the old SHA-1 hashes should have been preserved and stored in some format that would have allowed the retrieval of the new SHA-256 hash corresponding to an old SHA-1 hash.
I know this is just one tiny piece of the omnishambles, but: can't github sidestep the "you have to tell us what kind of hashes you have at repo creation time" by effectively deferring "git init" in practice until it sees what's being pushed? There may well be a good reason this wouldn't work...
It would make git compatible with newer projects that store SHA-256 hashes. For example Safecloud can be used for encrypted storage and streaming: https://safebots.github.io/Safecloud/
Ok, so the guy manages GitHub and doesn't want to pay the price.
Fair enough.
But the worrying part is that he seems really convinced that his security analysis of the vulnerability is complete, and that we should be convinced as well.
If you didn't need a strong hash, use a very fast one then, like MurMurHash.
It seems people are missing the point: it's not even the submodule incompatibility that's going to become an issue majorly (like python2 --> python3 but worse), the main issue is the loss of traceability for repos that changes in place (which I assume many will do). Imagine what will happen to these:
- SLSA and Provenance or SBOM data in the supply chain security that uses commit hash. All the previous images are now pointing to a non-existing commit
- All the documentation and tooling as the article calls out
- All your traceability links from your project tool to your git repo, they will lose all the past data as it will be dead links
So I hope there IS NOT a migration path for in-place replacement!
You'd think after all the failed (i.e. IPv6), and barely-succeeded (i.e. Python3) migrations the open-source world has seen, that the architects thereof would prioritise backwards-compatibility a little more highly...
schacon: Really like the "Independent Tree Hash Headers" idea.
How difficult would this be to get this functionality into git?
Would it cause any breaking changes with older versions?
Have you discussed this with any git devs to see if they are open to adding it?
Actually, this entire blog post came out of a short chat at Git Merge a few weeks ago with Jeff King. I argued more or less this and he didn't _entirely_ disagree, though he has good counterarguments on the list over the last few years, so I don't really know how he thinks about it ultimately.
I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"
Git might be the technology with the lowest understanding per user score.
Which makes it not surprising how major release of git flies under the radar of tech media coverage.
The first reason the author lists for why this will be bad is only an "issue" on Git hosts that don't allow repo creation on push (which is brain dead of GitHub). Any other host, you push your new repo, and it will see the hashing algorithm, and receive the contents accordingly.
Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a setting which can be changed.
I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.
Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
There are many disappointing errors in this article that one would expect an author with such prestige to know better.
Friends don't let friends build decentralized content addressible storage on insecure hash functions. Nor do friends try to convince you to use insecure hash functions. Nor do they appeal to authority. They also warn you against signing the output of insecure hashes. Finally they don't bother making argument that it's not worth upgrading today because we might have to upgrade tomorrow too.
I see someone has raised 17 million for a new ai git frontend.
oof. I can only hope that as many as possible git clients/wrappers that people actually use stick with SHA-1. Solving issues related to that not having happened sounds like the least meaningful work one could be doing...
Also GitHub should auto-detect the format on the very first “git push” and handle it then. If they actually require config to be set appropriately during the initial “new repository” dialog that’s just bad ux.
As for this being a looming disaster for industry… it’s inconvenient. And lots of things will break. But we will survive and come out the other side I’m certain. We converted from the Julian calendar to the Gregorian calendar 500 years ago. Surely we can handle this.
I'm going to completely ignore the first part of this post because I'm not interested in arguing about how severe the issues with SHA-1 are. I think it's accepted that there are flaws.
So given that, I'm more interested in the arguments for why migrating to SHA-256 is problematic.
The biggest issue I see, after skimming over it, is the submodule breakage for new projects trying to link to old projects. This seems solvable frankly, but is the only serious issue I see. Everything else will be worked out as software is updated IMO.
I think the only good argument here is that sha is maybe not a security control for git. I think every other argument in this article is incorrect
a) it's relatively fast and impossible in a practical sense for two different files to accidentally hash to the same value.
That is silly. We are not worried about accidentally triggering. We are worried about intentional triggers.
I dont know why people always bring this up for hashing. In any other context it would be considered silly. If someone said, the chance of triggering a buffer overflow by accident is low, we would call that silly as we aren't worried about accidental triggers.
b) second pre-image vs collision.
In a world of open source where we accept commits from randoms on the internet, i think collisions are just as relavent as second pre-image.
> b) so, you agree with the blog? "So, any realistic interesting attack vector therefore relies on a collision attack,"
My reading of the blog is that they are dismissive of collision attacks. In context of git, i disagree. I think there are plausible attack scenarios involving collisions, or at least, just as plausible as second pre-image.
If you mean do i agree with the blog that impossible attacks aren't possible? well yes obviously, but i think that goes without saying.
Then `git fsck` wouldn't be able to tell if an object uses the sha1 scheme or the sha1∘sha256 scheme, so it may need to compute the both hashes to check an object's integrity. Also, an object name alone wouldn't indicate that the weak scheme shouldn't be used to check integrity, so a malicious sha1 object could be swapped in in place of an sha1∘sha256 object (if second-preimage is found).
> In other words, hash collision attacks are maybe the dumbest possible way to get untrusted code on a system when unpaid open source maintainers and low-trust package management forges exist.
Yes. And we literally see this every other month with NPM
Thanks for writing this. I'd only been loosely following it and I hadn't realised how bad this is going to be. I have repos with tens of submodules and it's going to be a nightmare if any of them switch to sha256 in place. Not to mention I won't be able to use any new projects unless I rebuild my repo and all the submodules therein.
I thought the master to main thing was bad enough but this is going to suck. And just like the master rename it achieves basically nothing.
What is it about these projects that attracts people who just want to change things for the sake of it? Real engineering means coming up with a solution for backwards compatibility. This is just irresponsible and, frankly, a fuck you to everyone who will be affected by this.
« people who just want to change things for the sake of it » — Not for the sake of it, but to “make the world a better place.” The intention is noble (well, mostly, at least let's assume it is). Of course, there is disregard for history and her deplorables (or not so deplorables), comes with being “progressive”. Which is why I like Windows better (not 11 though, will probably have to go back to Linux at some point).
« What is it about these projects » — Maybe that they're “at the forefront.”
> I pull it from there because I trust that GitHub has its authentication game together enough that it's unlikely that anyone malicious pushed something there without the maintainer's knowledge.
The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.
This is very very easy to fix if you run into it.
1. Adopt git 3.0 if you can with sha256.
2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.
Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
> sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue
If you read the OP article, the entire point he's making is that this would never happen, because a hash algorithm being "broken" doesn't matter in practice, because true supply chain security has nothing to do with file hashes.
It's easy for _one person_ to fix. It's not easy for the entire git ecosystem as a whole. GitHub, large internal corporate git repos, CI/CD systems, projects with submodules, etc. The second half the article explains all of this.
It’s already broken, but even though it’s broken it’s hard to generate git collisions because of the repo metadata. It’s easy to generate (for instance) standalone PDFs with identical hashes, but doing this with git in a useful way is much harder.
That said, it’s still a good idea to migrate to a more robust hashing algorithm. Defense in depth, etc. Just because it’s a difficult migration doesn’t mean it shouldn’t be done.
You can certainly do this, as I said, this is Google's backup plan. But defaults matter. People will start running this and getting repos that are uselessly incompatible with other repos, tools, libraries and server instances. Having it as an option is one thing. Making it a default will cause a lot of pain for people who don't want to care about this.
No, my argument is that the change should not happen at all and nobody wants it and it gains the community very, very little but the default change is forcing it on everyone and most will be _entirely_ unaware - now having to solve problems that are difficult to understand. Defaults also matter when they are the wrong defaults.
Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
For the record, Y2K was not fud. It was very real, in a long list of datetime problems that are to come. Further datetime problems are coming at scheduled dates.
It was definitely FUD. There was a real problem (date counters would roll over), but the impacts of it were so ridiculously overstated that it eclipsed any sane discussion of the issue. We had people at the time predicting that planes would literally fall out of the sky when Y2k hit, which was never a realistic possibility.
Considering one of my team literally was carrying a Motorola satellite phone on the NYE/D of 2000-01-01 in case our mobile or land comms went down because the public transit was running all night, we took it very bloody seriously.
If there was an interruption to services on that night in particular, that's a serious public safety issue.
We had tested our system, and integration tested with all the other systems we connected to directly, but any FMECA analysis would show you that there were failure modes that we couldn't mitigate.
So people on planes falling out of the sky? Probably no.
People being crushed in a railway station on NYE? Possible yes.
That's the problem with deniers. When responsible persons take preemptive action to prevent tragedy, like with Y2K, the diners say it was FUD. When people don't take action, like with climate change, they say it wasn't important considering it's not them who's dead, totally discounting those who have suffered or died as a consequence. In summary, the deniers are so incompetent that they can't be trusted to correctly maintain a car, let alone civilization, considering they would never even the replace the necessary parts at the right schedules in their car.
As linked by another commenter in this thread, Linus worked out years ago that even if someone inserted a malicious object into the kernel repo, it would at best be a nuisance and not a major concern.
They address this very theoretical. In short: Yes. Which makes sense if you don't treat the hash as a form of security against malice, especially in the case of attacks that are already impractical, which is the entire thrust of the article.
Claude, make the hash use SHA-256 rather than SHA-1. No errors plsss.
Besides the possible implementation/deployment issues they will or will not face with this update, I can empathize with the idea that of not wanting to have a possible vector of attack in your system. Particularly today with AI being able to find novel exploits, I could see a future where a vulnerable hashing system leads to a malicious injection attack.
The author argues that "If I wanted to get untrusted code into Android, it's so much simpler to bribe or convince the maintainer of a popular downstream project" which is a really a red herring in this matter since that is literally a completely different issue that obviously no software update could ever fix.
Nonetheless I do agree with him in regards of how complicated and messy this whole process will be. Crypto migrations have been historically difficult, expensive and overall ugly, but not impossible...
This article is full of mistakes and misleading claims:
1) It's claiming SHA1 insecurity is theoretical, while SHAttered from 2017 was specifically a pratical proof of concept. The only reason Git wasn't affected, is because they didn't bother bruteforcing a git-blob prefix.
2) It's claiming collision attacks don't matter, only second-preimage attacks do. This is incorrect, collision attacks are enough for code-smuggling problems, when two repositories are on the same git commit (verified by the full commit hash), yet contain different code in their git checkout.
3) The Linus quote "The real security is in distribution" is arguing that "git's content-addressed system should not be used to address content". It's arguing that, in case of curl|sh, you shouldn't use a sha256sum-gate to pin the content to something you've reviewed, you should instead ensure curl is fetching from an https server.
1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
2) I specifically argue that even if both attacks were practical and cheap, it's still not the problem we should be focusing on.
3) Have you read this email (that I linked to)? It is almost the same general message (20 years ago) that this blog post is. It literally goes though a theoretical object replacement attack and how dumb this scenario is and so SHA-1 is fine.
https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...
I say it's impractical to exploit, which I think everyone agrees with.
Impractical for an individual, definitely. For a large org, maybe, but if the payoff was big enough? For a nation state level actor intent on doing something, absolutely not.
The go-to example is Stuxnet. Some countries wanted to attack Iran's nuclear enrichment programme, so they spent 5 years developing a worm that used multiple zero day exploits to attack a specific controller in a specific model of gas centrifuge. Could Mythos write Stuxnet? Unlikely, but a knowledgable team with access to it could probably write it in a lot less than 5 years.
'impractical' has very different values for different groups.
Yes, and in the days of major supply chain attacks and state-sponsored near catastrophes like xz, "it's probably good enough" starts to look incredibly naive.
Linus's "what matters is distribution" comment also doesn't make sense when merge effectively is distribute. Which, again, is the reality of supply chain.
Okay, fair enough. But does that make sense as a default setting then?
I can see that some things might have a risk profile that might possibly make all this costs still worth it, but does it make sense to have these unicorn projects effectively blow up 20 years of ecosystem?
Shouldn't the extra cost of doing something out of the ordinary be carried by whoever does something out of the ordinary?
This feels like a bridge to be crossed when one gets there (if at all).
__
FWIW, we actually do have a choice here. No one is forcing the industry at large to adopt an unpatched git 3.0 binary built from a source that makes that a default.
This should be a trivial overlay to carry around with effectively no downsides. So convincing whoever is steering that ship doesn't necessarily matter, as long as enough sane pragmatics agree on how defaults should actually be.
We are migrating all of PKI to the more costly and less efficient postquantum cryptography, even though nobody will reasonably use a quantum computer to snoop on your home IoT daily reports. I mean, I assume that what you are doing on your free time is not worth governmental attention.
The rationale of mass migration is that if you don't impose it, nobody migrates. This has notably been the case with famously insecure SSL parameters (512 bits RSA keys, PKCSv1.5...). And many companies may believe they are not critical, which might be true until it is not.
Case in point: you manufacture walkie talkies and suddenly your products have bombs inside. Or you maintain a compression library for free and suddenly you are shipping a backdoor to all Linux products.
> blow up 20 years of ecosystem?
What does that mean?
This is an apples to oranges comparison.
Stuxnet is essentially "boutique malware". You can buy it/have it built with enough money/resources.
Weaponizing a cryptographic algorithm with some theoretical vulnerabilities (but no by-design backdoor built-in) is a totally different game. And TFA is right that it's a dumb endeavour. You can probably "stuxnet" your way in for much cheaper.
the comment is highlighting the fact that git's use of sub-par cryptography makes "boutique supply-chain" attack easier to pull off.
> Impractical for an individual, definitely. For a large org, maybe, but if the payoff was big enough? For a nation state level actor intent on doing something, absolutely not.
In all the years since 2017, with all the orgs having huge GPU-filled data centers (and an interest in software security) has anyone demonstrated a real git collision?
Some systems have other properties that mitigate or prevent second preimage attacks - for example when you get an SSL certificate, CAs randomise the serial number. So an attacker can’t choose the checksum of the data the CA signs. Perhaps something in the design of git is similar?
'randomize' is possibly the wrong word, it could include some form of serialization/fingerprint.
And it’s also the question of different use cases.
Asking to trust in an authority (while the main authority Microsoft/GitHub has is essentially figuring out enterprise sales well enough to be acquired by a company desperately needing developers after fumbling badly in the 2000s) is exactly the opposite of my stance - it is a large corporation, with heavy employee rotation, with substantial exposure to various forms of regulatory pressure and to various forms of corruption.
Which is why the cryptography exists to prove the developer-to-consumer trust without trusting the intermediaries. Yes, I do check GPG signatures. Yes, I do include a git commit hash in the binaries I build. And surely I want to make sure that this doesn’t mutate because some unknown engineer at GitHub had a bad case of gambling debt.
Impractical for everyone. Because if you want to do this, there are better attack vectors, even if second preimage was easy, even though it is impossible.
> 1) I link to the SHAttered paper, as well as Shambles. Git projects were not affected because it is an inefficient attack vector. I say it's impractical to exploit, which I think everyone agrees with.
It seems unlikely it will stay that way forever. Typically attacks get more efficient over time as researchers find improvements, not to mention computers getting better.
In 2015 it was estimated to cost $100,000, now the estimate is down to $10,000. Where will it be in 2035?
I specifically argue that it doesn't matter if it's $1 and base my argument and solution around that. So it's irrelevant where it is in 2035.
You still mentioned the (1) part, which is what i object to. I agree its not fatal to your argument.
So you’d rather have a weak default for everyone rather than updating it? Share some PII then, I’m sure a mod will block it.
Even if it costs zero to create. How do You force people to pull from Your repo?
The XZ incident would have been so much worse if it would have involved a collision, pushing one object variant to github.com, and one variant to git.tukaani.org.
Then you would have security researchers making conflicting claims depending on which repository they first pulled from, even though they are on the same git commit hash.
In GitHub, fork networks are represented as one repo on disk (this is the main optimization that makes community pull requests possible at all). So you could fork a repo and then push an object with a hash collision to it to change a file in the original, in theory.
(In practice this is harder, as the article mentions, because the new forged object would have to be a valid gzipped git object of the same length. And GitHub probably knows about this type of attack and might just, for example, prevent existing objects from being overwritten)
Hack into the system holding a trusted repository, and swap in your variant with same signature?
Social engineering?
DNS cache poisoning?
Basing your cryptographic advice on a 20 year old opinion-piece from somebody with no background in cryptography is not the flex you think it is.
Appealing to lack-of-authority without actually explaining in what way his argument is wrong is significantly worse.
It's not cryptographic advice from an opinion piece, it's a statement about intent from the creator of the software in question.
But the Linus piece is sound
... for Linux
... and developers working for it constantly
the attack wouldn't work. Joe Schmoe? It's worse than just "being compromised"
You have repo of dependency locally, let's assume you downloaded good copy, the commits get compromised, you're safe.... right ?
Nope, if there is build server along the way and ESPECIALLY if it practices building from clean state every time, the build might be infected while your local copy is clean, giving no chance to notice it, unless your entire chain including local builds are reproductible AND you actually check it
In agreement that this is good old fashioned cargo-cult security theatre, but tom7 also coined a more catchy phrase for this, he calls it "toxic max-security." http://tom7.org/httpv/httpv.pdf
Oh I call those people „TLS antivaxxers”.
Conveniently Tom didn’t mention anything about Edward Snowden and what he published. That was basically start of TLS everywhere.
Then he didn’t mention ISP idiots that were actually injecting ads to cute websites like Tom’s. I hope Tom likes when his website is used by ISP to make money on ads he doesn’t have any control over.
Then he goes on to criticize certificate transparency, but it works. Companies got kicked out from trusted root program because they were doing stupid stuff like making certs they shouldn’t.
Let’s not forget glorious state of Kazakhstan where without TLS they would just listen to all traffic - well with TLS they were trying to pull MITM but were uncovered and got their stuff removed by TLS ecosystem.
> I hope Tom likes when his website is used by ISP to make money on ads he doesn’t have any control over.
You mean like how Google makes money showing ads against your content that you don't control? You mean how basically every ad network works?
If I recall, tom7 was at odds with chrome throwing up a warning at users trying to visit his website because he didn't support TLS on it. He wasn't against TLS. Calling them a TLS antivaxxer is not accurate.
You are basically resolving a non-existent security problem by generating a far bigger security problem, because I'm 100% sure that a ton of software just assumes that a git commit hash fits in a `char[40]` and thus will buffer overflow like hell if they try to operate on new repositories.
And we are talking about who knows how many tools that work with git built in the years, and this is also made it worse from the fact that most tools just invoke the git binary and capture its output instead of passing from a library.
I like more the solution proposed at the end of the article, do not change sha-1 but instead, if you are relying on git commit for security purposes (that was never the intended use) add another header to the git object with a sha-256, so that with the small expense of computing the hash twice you don't break 20 years of existing tools that make the assumption of the git commit being 40 character long.
I would have liked sha-256 truncated to 40 characters (by analogy to sha-512/256 that'd be sha-256/160). Cryptographically that's probably fine. But 160 bits is not a lot. And probably fine doesn't tend to inspire confidence in the field of cryptography. It's not a very well studied scheme
An advantage of a hash with a different length is that a full length sha1 commit hash and a full-length sha256 commit hash can't be confused for each other
Except that git commands generally accept partial hashes...
if you have buffer overflows today, you have buffers overflow.
git doesn't change that.
everyone is already using sha256 everywhere. i am. sha1 is only still around because github forces it for pretty urls
Persistent problems with increased hash lengths seem unlikely: if bad git clients have erroneous truncations or buffer overflows, they can either fix them instantly, problem solved (they have had many years to implement and test SHA-256), or fail to fix them and be written off as obsolete and incompatible: problem equally solved.
The problem of a SH1 collision happening by coincidence is vanishingly low and theoretical.
Nothing else matters.
Git hashes are not supposed to be a security mechanism. If your basis for trusting that you have the right checkout is the git hash, in a situation where you have legitimate concern about untrusted parties manipulating remote repositories, then you're simply wrong.
Sorry, no.
When I check out code from a git repository in a pipeline using a git hash, I expect the code to be exactly what has been reviewed by me under that hash.
Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?
Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Oh, that would never be a problem for widely disseminated, popular, open source project, so it doesn't matter.
Dealing with potentially-hostile hosts is quite common, actually. See for example how most Linux mirrors work, or Subresource Integrity with HTML.
Turns out securing a service to transfer a single hash is a lot easier than securing a service to transfer gigabytes of data.
Even if I don't fully trust Github, it is still incredibly convenient to be able to upload my code there and then send someone an email telling them to fetch commit `123abc` from some repo link. As long as my email isn't compromised, that should be secure.
Rumor has it that GitHub has a flat namespace for commits. They don't store "user1/repo1/abcd1234" in one file and "user2/repo2/abcd1234" in another. Both references point to the same commit in a global shared space. If the hashes are truly unique, then that never matters, because the odds are approximately 0.000000000...000 of you and I accidentally generating the same commit. However, if I see that you pushed commit abcd134, and then I can build and push a colliding commit, and the backend doesn't check uniqueness before writes because the odds are infinitesimal that it'd ever matter, than voila, I've updated your repo by writing to my own.
Or if first writer wins, and I know that you have a popular non-GitHub repo that you're about to migrate into it, then I could pre-poison the namespace by writing my own version of a commit that I see you already have in Codeberg or Savannah or wherever.
I don't swear that this is how GitHub actually works, but I've had knowledgeable friends swear up and down that it is. And honestly, it'd make sense. They could shard storage by the first 4 digits of the hash or something, and that'd be vastly more efficient if all commits were writing to the same space.
> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?
Yes, they ought to be good enough. That has always been git's security model.
Note that the commit doesn't have to be communicated over the same channel as the git data.
> Well, what about someone who is fetching the commit from that server for the first time and has nothing to compare the hash against?
Then they're vulnerable. But what about somebody learning about the trusted hash in another way, e.g. a build server getting an internal call authenticated by an authorized developer?
Just because you can think of an insecure way to use git hashes doesn't mean there aren't any other, secure ones.
> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries, you think that git hashes are good enough?
How does this matter? When I have a machine that I trust and a git hash that I trust, I don't need to rely on the transport mechanism. As long as the content cannot be forged to match the hash, the transport is completely irrelevant.
> And so if you don't trust the server that is hosted on or the security of the transport mechanism like TLS/SSL, such that the content may be manipulated by adversaries
Wait, why is anyone expecting that to be a good idea at all?
I mean, Git commit signing should be used more often... then you can actually trust the person signing, not the distribution method
But you still need SHA256 for that
Sorry, no. For trusting code there is code signing. A sha-1 hash is cryptographically the wrong approach for it or we wouldn't have RSA and ECDSA algos.
And what do these cryptographic signing algorithms do? They sign a cryptographic hash of the data … Which the SHA-family of hashes are.
If you rely on the commit SHA-1 as a integrity verification it's your problem. Git was never intended to be used as an integrity check.
Git was maybe never intended to be used as an integrity check, but SHA hashes definitely were and are.
And if git provides a cryptographic hash over the content, I don't see why it shouldn't be used to verify the integrity of the checked-out content.
I think perhaps you used the wrong words there. The hash is an integrity check but it was never intended as a security check.
>I expect the code to be exactly what has been reviewed by me under that hash
Sorry, no.
Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
> Everything else would just be a crazy invitation to make supply chain attacks uncircumventable.
What? Uncircumventable? Logically equivalent, I read your statement as "So if all cars are not blue then they must be red"? This does not follow...
I like the last part of the article, which proposes a (very reasonable sounding) extension for people who care (more) about their code-sec. You should be able to swap out your hash algo without having to rebuild your content addressing system.
Separation of Concerns, people...
> Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
This doesn't seem to be a useful definition. Would you classify every computable algorithm as insecure, because by generating a random bitstring, there is a (very, very low) probability of guessing the hash/secret key/solution/signature?
> Again, hashes are by definition, insecure. I.e. you don't have a guarantee that the hash references the same commit, just a (very, very strong) probability that it does.
Cool, you've just defined the foundation of signatures and web encryption as insecure. What next?
Also, you can only be so certain about any piece of data no matter what you do. With a non-broken hash you can make the collision chance be a trillion times lower than the chance you're hashing the wrong data to begin with. That's as good as gold, well actually better than gold.
Sorry, but aren't there signing and attestation mechanisms built into git already? Why not use those?
> Git hashes are not supposed to be a security mechanism
Probably a naive question, but why not kill two birds with one stone if it can be done for a reasonable cost?
Because you're not killling two birds; you're not killing the security bird with a better content hash.
A SHA-256 sum, though very good, only assures you with great confidence that you're looking at the same thing you looked at before, or that someone else is looking at elsewhere.
It is not a digital signature, and we don't want digital signatures to serve the role of content hashes.
Speaking of signatures, we have support for them in Git; you can use gpg to sign commits, and set it up to be done automatically.
Nobody is going to fake your commit such that the fake has the same SH-1 hash and your GPG signature.
The worry there is that the key holder (whether the legitimate one, or a malicious party who got a hold of the key) somehow does this: creates a new commit, signed with their key, which somehow has the same SH-1 as an existing signed commit. The git hash includes the GPG signature, so there is a significant layer of difficulty there which is likely harder than faking an unsigned SHA-256 commit.
Dude... please bow out gracefully...
The attack is I pre-author `Makefile => foo: echo "hello"; bar: echo "world"` along with `Makefile => foo: echo "hello"; bar: rm -rf / ; /* $ELDRITCH_SHA1_SPIRITS_GO_HERE */` that both hash to `ff1234...`
I then prepopulate the repo with `echo "hello"`, wait 6-9 months, then submit a commit for `echo "hello" ; echo "world"` and keep (in my back pocket) the alternate implementation that also includes $ELDRITCH_SPIRITS to force a collision and MY predetermined change in functionality.
I then have free choice as to whether I serve them "hello world" or "hello && rm -rf", and THAT's the plausible problem to avoid: the ability to "cloak" content anywhere within the repo if you have enough $ELDRITCH_SPIRITS and GPU's.
You have _really_ good points, but are woefully confused. The proper answer is (would have been) to include `tree ff12354...` along with `tree-sha256 abc123456789...` for another 20 years along with a `[git.hash_strictness]: default/lazy/strict`, and some oddball `git-rerere` type packfile extension which lets you map `sha1:ff1234... => sha256:abc123456789...` "transparently" rather than the horrific situation you're laying out (correctly!) that forks the ecosystem in to "longhash" and "shorthash" when most repos don't even care in the end.
Please educate yourself what a merkle tree is. It's a well understood building block of various security systems, including certificate transparency (which explicitly uses sha256).
You refer to PGP signed Git objects, but you also argue:
> Git hashes are not supposed to be a security mechanism
Guess what the Git PGP signature is signing.
You do understand that most code signing is just PKI over the top of SHA hashes, right?
The only substantial difference is that the author vouched for these particular snapshots of code, in a way where nobody else can intercept the communication and substitute a completely different hash than one the author previously signed.
In fact if you can find a SHA collision, you can peal the signature off of the legitimate payload and slap it on the colliding one.
> Git hashes are not supposed to be a security mechanism.
That you are wrong. Many people and orgs rely on this.
Does not matter if you think it's stupid, for them it is. live with it.
Exactly! The recommended way to use GitHub actions is to use their deploy hash as the version to prevent hijackings. It was never the intention (nor would this have been envisioned when git was created), but security needs to be applied to the way tech is used.
Supposed to be or not, it is.
Package managers even use it. E.g. you can have a cargo dependency pointing to GitHub at a specific commit. It's definitely intended to provide end to end security without depending on GitHub being secure.
Also git submodules.
And I should add: not just github compromise, but supply chain / original author replacing the contents.
Absolutely the SHA-1 is treated as "authenticating". Cargo.lock (for regular crates.io dependencies) are confirmed using SHA-256.
> Git hashes are not supposed to be a security mechanism.
Commit signing indicates otherwise.
Linus said it was a security feature in that big presentation he did about Git.
> Git hashes are not supposed to be a security mechanism.
They absolutely are, both when you're using signed commits (i.e. GPG, SSH, and S/MIME) or just identifying a given commit/repository state by hash via a secure channel and then serving object data over untrusted transports.
It's entirely possible that you don't use either, but that's certainly not true for everybody.
Calling anyone using either to be "doing it wrong" is borderline gaslighting: Git used to have these security guarantees, and just because they're now broken doesn't mean they were never there in the first place, or that it was stupid to rely on them.
That's what a lot of the responses are missing, which is why the response to them in turn is both no and yes. There's no published threat model that I know of for the properties that the hashes are supposed to be providing for git, which means anyone can make up any required property they like and then confidently state that SHA-1 provides or does not provide it, see this discussion for examples. The result is, to use another Linus quote as the OP has already quoted him in the writeup, "people wanking around with their opinions".
We use SHA-1 in our storage mechanism, which predates git. There is a (quite long) written threat model. Someone being able to generate collisions with an enormous amount of effort under just the right conditions is not a threat under that model.
No, that argument misses the point. Cryptographic hashes enable you to trust, that I only have to review the changes since the last trusted commit. That is more important for the developers and maintainers of the software and less for the users.
In order to work for the users, a thing first has to work for its makers.
There are a lot of tools where the lines blur between “for the users” and “for the development team” because the users benefit from some things that make the developers’ lives easier.
You could pull a malicious colliding PR and reject it. Then you pull the other half of the collision without realising it is, and it's something good and you merge it. But your CI server already saw the malicious one and thinks it's the same, so builds the malicious code
The important point is that the switch is breaking backwards compatibility. The proposed solution, a new independent hash just for verification makes sense.
This is the top comment and yet says absolutely nothing to refute anything in the article, merely declaring that it's "full of mistakes". I feel like people haven't read the article, or this comment, and are essentially religiously predisposed to so-called progress, no matter the cost.
Your points are full of mistakes and misleading claims.
I don't simply mean to be disparaging - its important to security that the people making the decisions are a) competent b) can read, otherwise any "security" decisions they are making are at best probably insecure, and at worst, causing active harm and insecurity, DOS, etc...
Generally, I think you didn't read (or at least comprehend) the article:
1) Your assertion is factually incorrect. Practical proof of concepts are not "CVEs exploited in the wild". The article is not claiming that SHAttered is not correct, it in fact references it.
2) I don't think you read the article. The article claims the exact opposite, and in fact addresses the issue WRT to distribution.
3) Your sentence here is very confusing - I'm going to give you the benefit of the doubt and presume that you mean to say that the problem here is in the security of the naming. HTTPS itself has nothing to do with distribution security, that would be DNS/SecDNS (IFF you are using git+http protocol, then HTTPS is relevant, but not to git otherwise). But this is exactly what the article was talking about, the distribution is the security, not the hashing algorithm.
Also talking about quantum computers breaking SHA-256 means he knows a lot more than the rest of us or not nearly enough to pipe up about the subject.
For reference: https://git-scm.com/docs/hash-function-transition
Notably, a few of the featured author's reservations appear to be addressed. According to the Git docs:
- Objects can be referred to by their old, SHA-1 name or their new, SHA-256 name. This means old refs in docs and comments and such remain valid. The mapping between SHA-1 representations and SHA-256 representations appears to be intentionally bijective a.k.a. 1-to-1 (assuming no hash collisions), so that it could be re-computed on demand. The constraint of bijectivity appears to be the source of some limitations, ex. no mixed repos and submodules needing to match hash algroithm, but also bijectivity has strong benefits like the following items.
- A bi-directional dictionary is maitained from SHA-1 to SHA-256 names so translations between the two don't required re-hashing objects. This table could be recomputed on demand due to the bijection between names; it's only a performance optimization.
- A local SHA-256 converted repo (including an SHA-256 converted submodule) can interoperate with an SHA-1 only remote transparently to the remote server by translating names using the lookup table.
- SHA-1 based GPG signatures will be preserved. A commit can be signed based on its SHA-1 representation, its SHA-256 representation, both, or neither. The bijection means the two types of signatures are in a sense interchangeable, or in other words the bijection between object representations implies an equivalence relation on signatures. An SHA-256 converted repo can quickly validate an SHA-1 based gpg signature using the lookup table.
If we are changing to sha256 because sha-1 is broken, how can we trust all those mappings and signatures?? This makes it even more confusing as to why this change is being made
Git moved to a modified SHA-1 (let's call it SHA-1′) back in 2017 which prevents the particular attack described in SHAttered, so the SHA-1′ object names and signatures can be trusted for now. But the known collision attack on the underlying algorithm makes it more likely for SHA-1′ to be broken in the future too. Researchers could find variations or improvements of the weaknesses used in SHAttered to attack SHA-1′ as well. In other words, there are no known practical collision attacks, but there are known avenues of reasearch that could lead to practical collision or second-preimage attacks being discovered.
The idea, then, is to migrate all git history to a stronger hash before that happens. If we waited until an attack was found, it'd be too late; all git history would be suspect forever into the future (disregarding extensive auditing), even if a stronger hash was used retroactively. Full SHA-256 doesn't have known practical collision attacks, and it doesn't even have known research avenues likely to lead to practical collision attacks, so is considered more future-proof than SHA-1′.
This idea of future-proofing is common in cryptography. For example RSA-1024 keys are now considered practical to factorize (with a supercomputing cluster and a few months), but that's okay, because, foreseeing the possibility, "everyone" switched to RSA-2048 or stronger a decade ago. Now we're adopting new key algorithms to protect against possible quantum computing attacks in the future.
Yes, I understand all that. The original author points out that completely ditching sha-1 will be a huge pain, very costly, and won't meaningfully improve security. You responded by saying we aren't completely ditching sha-1, which seems to lend credence to the argument that sha-1 is still safe enough to use. That's where I'm confused. Having two hashes for each object in a git repo, a sha-1 and a sha256, is not the same as wholesale deleting your rsa-1024 keys and replacing them with rsa-2048. If I'm only looking at the sha-1 and all my tools only look at that, then the sha256 isn't helping me. Is it?
Ah, the plan is to eventually completely ditch SHA-1, but only after "everyone" has already converted to SHA-256, which is hopefully before SHA-1′ is broken. See sections Object names on the command line and Transition plan from the link in my top-level comment.
If you trust the sha-256, you know the commits you have and their ancestors are not evil. So, the sha-1s for those commits can be trusted.
You should also check the mapping table's sha-256 hash, to know that it's not evil.
If you later get an evil commit trying to masquerade as one of those, Git will presumably notice that there are two commits with the same sha1 in the same repo, and bail loudly.
I wish I could upvote this more. The reality is dwarfed by a mega thread commenting on noise.
I threw in an upvote, there is previous history in software for version migration. With version 2 to 3 being the migration here, surely there's a way to stuff 2 bits of info somewhere in a git dir and just patch the version 2 today to start looking for those bits.
Seems like the first rule of HN
now I'm confused: if sha-256 is only used for new repos (default) and sha-1 continues to work for sha-1 based remotes, so a git clone, git push just keeps sha-1, then what exactly is the remaining issue? to me this sounds like a perfect and user friendly migration strategy.
The third-party tooling concerns seem plausible. For example, it seems like the intension for Git hosting is:
1. Git users all transition to SHA-256 while git hosts stay on SHA-1.
2. Once almost all Git users are on SHA-256, Git hosts flip a switch and convert all hosted repos to SHA-256. Git users on SHA-256 don't notice anything, because the only thing that changes is how the client and server negotiate which objects to send.
3. Git hosts disable SHA-1 support. Git users on ancient clients need to upgrade or switch hosts.
Instead, by the author's account, Git hosts are deciding to convert repos to SHA-256 individually and placing the decision of when to do each conversion upon their users, who may not be informed about the situation. To me, that seems like the wrong decision, but then I don't run a Git hosting service.
No doubt other tools will take some time to catch up too. For example, libgit2 is perhaps the most popular third-party Git client, and it still doesn't support newer Git features like reftables (circa 2020). Though libgit2 does appear to implement SHA-256.
One of my favorite fun facts about Fossil SCM (another source control by the devs of sqlite) is that they patched their use of SHA1 6 days after the shattered attack was published:
"Both Fossil and Git started out using only SHA1 hashes. But when the SHAttered attack against SHA1 was published on 2017-02-23, the need to migrate to a stronger hash algorithm was recognized. Fossil added the ability to use SHA3-256 as an alternative on 2017-03-01 (six days after the SHAttered attack was first published). SHA3-256 is now the default for all new repositories and check-ins in Fossil, though older check-ins that occurred prior to SHAttered can still use their original SHA1 hash. Hence, no repositories had to be rebuilt and no hyperlinks were broken."
https://fossil-scm.org/home/doc/trunk/www/hundredandone.md
To me it's so interesting watching in realtime Git is still battling with this decision and for Fossil it was just another week of development.
That whole page is fun to read. Another fun fact somewhere else in the docs is that Fossil uses a grow-only set to store commits. They came up with this scheme some years before it was formalized by CRDTs!
It's easier to make world breaking changes when the world is really small. If git could magically just get everything and everyone to cut over and use git 3.0 in a magic instant, it wouldn't be having this problem.
It seems Fossil can use multiple hash algorithm within the same repository. Yes, that decreases overall security since the hashes pointing to older objects can still be spoofed.
Here's a stoichiometric bird for you:
:%s/git 3\.0/python 3.0/g
Or IPv6 or windows 11.
:x
Neovim's lazyvim plugin sucks because it takes over H C and L.
Huh wha? What problems are solved by getting all Windows users onto Windows 11 in particular?
I mean, there are two things here. One is how difficult it is to have a different hashing mechanism. Brian and other heroes in the Git core group have done amazing work to make this _technically_ possible on a repo level. To test some of my theories, I trivially implemented MD5 and an insanely dumb and easily breakable hash backend. It's not _hard_ to change the mechanism now. It's about the community.
Fossil isn't difficult to change not because it's technically harder for Git but because Git has a community and ecosystem that Fossil does not. The cost is not in the individual project for Git, the cost is because there is _so much_ in Git and this bifurcates everything.
Also, interestingly, Git today does _not_ use a straight SHA1 because of these attacks. It uses `sha1dc`, a slower collision detecting variant that specifically checks for this vector of attacks. So currently, Git's SHA-1 variant is not susceptible to the SHAttered/Shambles attacks.
Sure but that's just kicking the can down the street. It will eventually be susceptible to a collision attack. As will SHA256 btw. What then? Are we doing to go through all this again?
Online video has handled this. There are various codex, container formats and transport protocols. The TLS handshake does this. The ability to deprecate and replacing the hashing algorithm should've been built in from day 1.
As an expansion on that, as far as I know there's only one implementation of Fossil. There was at least a Java one for git, I believe, and I don't think upstream git is related to libgit2 either. And GitHub is doing… something not visible from the outside.
(… looking at the parent, though, I imagine there might be some information from the inside…)
> As an expansion on that, as far as I know there's only one implementation of Fossil.
That's is, since only recently, no longer strictly true: the age-old libfossil recently got client sync support, so is now (aside from _serving_ repos) essentially a standalone impl (its own developer still uses fossil(1) stash, patch, and diff -tk features, but otherwise uses libfossil's counterparts).
Also, Dan Mestas has <https://github.com/danmestas/go-libfossil>, a Go library for working and fossil, and he is also working on <https://zeitforge.app/>, a clean-room impl. of fossil (whereas libfossil is largely ported directly from fossil(1)) which even goes so far as to _not_ use an sqlite database for its file storage.
Dan Mestas and Dan Shearer are working on finalizing RFCs for fossil's sync protocol and artifact format, and zeitforge is created by carefully managing LLMs which are reading that draft (but not the source code of libfossil or fossil).
I personally have been using `go-libfossil` in a project! It's got some sharp edges but insanely useful in avoiding CGO and using fossil at the same time
Agreed! This isn't a tech dig at all.
To me it's more of a reality of creating a tool with a huge active community and a community of contributors and creating a tool with a small team and small community.
in fairness, git switched to sha1-DC in may 2017, so they were only a few months late in mitigating it.
Isn't sha1-dc just checking for a small list of hard coded disturbance vectors? Meaning that script kiddies can't easily reuse exact values that researchers spent a bunch of time precomputing. That's far from being a good long term solution.
I'm not an expert on this but my impression is its a bit in the middle. It isn't a great long term solution, but finding new disturbance vectors is very hard. Its not like someone can just go find more on a whim.
Yes.
That's impressive. I suppose they had a more flexible architecture to make that change so fast.
Is there any writeup on why it was easy for them and not for git?
My guess is that it's less about the architecture and more about the blast radius and the number of users
no it's about the architecture and design: fossil enables coexistence of both hashes in one repo
https://fossil-scm.org/home/doc/trunk/www/hundredandone.md
27. Fossil allows both legacy SHA1 hashes and newer SHA3-256 hashes in the same repository.
Hmm probably nothing architecture wise. It's probably just the fact that fossil is developed by way fewer devs.
From the skim I read of this article it seems both projects arrived at the same solution: support both but make SHA-256 the default.
No it’s different. Git supports both but you cannot mix them in a repo. Fossil allows mixing them in a single repo. This is clearly stated in the link.
They should let you mix it in a repo then.
SHA3-256 != SHA-256
Oops! I admit I don't have much knowledge on hashing algos. I thought it was a shorthand
I always liked the ipfs concept of a multihash where a hash algorithm id is stored with the hash, I don't know how well it worked in practice, but in theory all applications of the protocol are now forced into a world where there are multiple hash formats and it can and will change.
We don't need that. Introducing an unknown hash algorithm itself is a security issue.
This problem isn't hard or new. Just look at things like a TLS handshake. You need to separate the protocol from the storage implementation.
I personally believe the initial Git programmers were too in love with the efficiency of doing a bitwise 160 bit comparison on the stack and they sacrificed the known issue of changing the algorithm to do it. A decade earlier the same thing had happened with MD5.
The saga continues with that proposed terrible github UI in the article (from a github cofounder!)
Feels like Leadership of Fossil SCM is better than that for git.
(Because: This very old problem is ongoing with git, and is closed with Fossil.)
can I have your email to get in touch
Linus Torvalds in 2007:
> but the point is the SHA-1, as far as Git is concerned, isn't even a security feature. It's purely a consistency check. The security parts are elsewhere, so a lot of people assume that since Git uses SHA-1 and SHA-1 is used for cryptographically secure stuff, they think that, Okay, it's a huge security feature. It has nothing at all to do with security, it's just the best hash you can get. ... [1]
[1] https://www.youtube.com/watch?v=4XpnKHJAok8&t=56m20s
So Torvalds used SHA-1 purely because he needed a hash function with no other property than identifying content.
It's of course possible that SHA-1 was not originally intended as a security feature, but according to Hyrum's law, every observable API behavior becomes something that somebody starts to depend on, so if you start very publicly shipping a cryptographically secure hash, you better keep it cryptographically secure.
At the very latest, this fact was cemented when first-party git commit signatures started depending on the security properties of SHA-1.
> if you start very publicly shipping a cryptographically secure hash, you better keep it cryptographically secure
So if my API happens to return text strings that always happen to have an even number of characters, I better make sure that all future versions of it also always return an even number of characters, just in case some moron decided to bank their application's functionality on that? No. If you decide to write a fragile application tethered to some incidental property of some upstream software, your application deserves to break.
> No. If you decide to write a fragile application tethered to some incidental property of some upstream software, your application deserves to break.
Just because you never made any promise regarding one aspect of your API doesn't mean that you're absolved from responsibility when you choose to change it. If you know for a fact that many users rely on it and you choose to break it, you need a good reason. That kind of balancing act is part of your job. If you don't respect your users, perhaps development wasn't the right career choice.
As developers we're of course constantly tempted to rename things that we named poorly, or change a schema that is no longer optimal. But we must always take a step back and think about the downstream impact.
> I better make sure that all future versions of it also always return an even number of characters, just in case some moron decided to bank their application's functionality on that?
Oh yes, this does happen. There's even a name for that: ossification (https://en.wikipedia.org/wiki/Protocol_ossification). You can't change your API/protocol, because "some moron" started depending on implementation details.
It's easy to say "your application deserves to break" from an ivory tower, but it's often not easy or viable to fix it (for instance, it might have different owners, it might no longer be maintained, it might be more expensive to change, etc). And the change which broke that application was not in it; the blame naturally goes to what was changed last.
Hey, I'm just the messenger here, if you don't like it, take it up with Hyrum ;)
But seriously: If you can afford to break your user's applications if they "deserve to be broken", sure. Many API maintainers can't, or at least don't want to.
In the latter case (which is honestly the norm rather than the exception, at least for public APIs), yes, you should better think about all implicit API contracts your API shape might be projecting. That guideline has served me very well through my career, at least.
> It's purely a consistency check. The security parts are elsewhere
Sorry if the video answers this, but how does commit signing work if it doesn't rely on the hash algorithm being resistant to at least second-preimage attacks?
Things changed since 2007
Thanks. (More detail: Git added signed commits in Git 1.7.9, which was released in 2012.)
How they changed?
For all practical purpose SHA-1 is a bad hash function, it's slow, it's insecure.
If SHA-1 is not a security measure, why do you even sign the commit. It doesn't make any sense. You give a strong signature on something weak.
Buddy, you are naming the things that changed.
Well they didn't sign commits in 2007, and there wasn't github in 2007, so it wasn't a security measure in 2007.
Exactly. But that's why I think SHA was a mistake. He should have gone with something like murmur to avoid all this frothing at the mouth.
It was a mistake to assume a fixed algorithm in the repository format and client-server protocol. I remember being surprised when I learned about that choice, being familiar with cryptographic protocols and formats where the hash algorithm is usually a parameter that can vary for each concrete hash.
> being familiar with cryptographic protocols and formats where the hash algorithm is usually a parameter that can vary for each concrete hash.
This flexibility ("agility") in cryptographic protocols is often seen as a mistake today, actually.
> It was a mistake to assume a fixed algorithm in the repository format and client-server protocol.
See also perhaps Wireguard, which touts itself as not having "cryptographic agility" because they wanted to avoid all (perceived) problems and complications of IPsec. But now that PQC is (allegedly) approaching there's no easy to update things because (AIUI) there's no negotiation possible in the protocol; you're basically standing up a 'Wireguard 2.0' that runs separately than the original.
1. As others have noted, flexibility in cryptographic protocols is generally a mistake.
2. The hash function is not used for cryptographic purposes!
Linus, in 2005, couldn’t have gone for murmur, from 2008.
Was there “something like murmur” in 2005 that’s cryptographically better than SHA1?
Yeah, in 2005, SHA-1 was just about the best you can do given the constraints of the time (without picking something much more esoteric, much slower, etc.). Using SHA-256 at that time would have been noticeably slower on the computers of the time, and made repo metadata take up a lot more space.
If that was true, i doubt git would have switched to using the slower version of sha-1 that detects attacks.
didn't he create git in like a weekend?
[dead]
Frothers gonna froth, though.
Hmm, I swear he disagreed with this in a different talk. But maybe I’ve conflated a presentation by someone else with one of his.
Hard coding a specific hash algorithm was a big mistake.
Have you seen algorithm-agnostic protocols, like TLS v1.2 and IPsec? It always turns out to be a bad idea. Future upgradability is fine but you really don't want implementations of the protocol to be incompatible with each other and you don't want attackers to have any chance of tricking you to using insecure protocols.
Flexibility made sense in 1995 when nobody was sure which algorithms would stand the test of time. Even in 2005 it was unnecessary and in 2015 it was an outright liability. If you have a good algorithm just specify the good algorithm, don't let the parties negotiate either a good one or a bad one.
Your argument is weird because it assumes we can know if a hash function will be secure forever. Assuming git would use SHA1 forever seems shortsighted.
That was true in 2007, but it stopped being true once people started signing commits and tags. A GPG/SSH signature on a commit covers the commit object, which names its tree and parents by SHA-1, so the signature is exactly as strong as SHA-1's collision resistance. If someone can prepare two trees with the same hash, a signed release tag vouches for both of them.
SHA-1DC blocks the known SHAttered/Shambles-style attacks, which is a good stopgap, but it detects known techniques; it isn't a hash you can reason about.
Code signing went through the same thing. Authenticode signatures with SHA-1 digests are effectively distrusted on Windows now, and that migration hurt precisely because it was put off until it was urgent.
Linus is one of the last remaining champions of rationality in large-scale software projects. Otherwise, it’s filled with devs who love to leave their brains out when it comes to practical scenarios. Anyone who thinks there is a security issue here is an absolute moron.
Rationality is the key point. A lot of devs approach their work in ways that are not rational or practical. Like the dev who wants to build a gold filigree decorated elevator that can handle ten thousand pounds of cargo and moves with the smoothness of a magnetic levitation rail in order to reach the second floor when all you really need is a ladder for the 2 people that will need access.
I don't understand why Git is not making the SHA-1 and SHA-256 modes far more compatible with each other.
SHA1-hashed objects should be able to refer to SHA-256-hashed objects, although this seems somewhat pointless.
But SHA-256-hashed objects should also be able to refer to SHA1-hashed objects, with a major caveat: if those objects themselves are part of a collision pair, then there is a genuine problem. But this is avoidable! Suppose that Linux decided to migrate to SHA-256. The upstream project could choose a pair of dates, say January 1 2027 and March 1 2027. Up to the first date, maintainers would be welcome to submit hashes of objects that are not yet in the repo but that they think they might submit later on, and, on that date, the upstream tree would finalize the list of these objects and reference it in the repo (with a new mechanism for this purpose). Effective the second date, the repo would start publishing SHA-256 commits and would never again accept a SHA1-hashed object that was not in the repo at the cutoff date or referenced as part of the Jan 1 block.
And now it would be impossible to get a new SHA1 collision in to the repo.
The only new git features needed would be:
a) actual compatibility so that a SHA-256-hashed object could reference a SHA1-hashed object
b) a new object type that's a list of allowed SHA1 hashes (or probably a tree of them) that is itself hashed with SHA-256 and a mechanism to link to one of these from a commit
c) a policy mechanism to set a repo to only allow SHA1-hashed-objects that a reachable from a preconfigured SHA-256-hashed commit
Emily's talk does a pretty good job of summarizing the issues with intermixing the hashes: https://youtu.be/eJJp0RE7cd4
They begin by saying mixing can never work but don't address the option of dual hashes... Then the interop plan is basically a secret git format that maintains both sha1 and sha256 hashes FOREVER.
I don't see why all users wouldn't want to keep both hashes.
I skimmed the video, and I didn't quite catch that. Near the end of the video, however, she did mention that interop is in the works[0].
In any case, even if Git 3.0 were completely incompatible, it would suck, but it's not the end of the world. You just treat it as if you were migrating from one SCM system to another. CVS -> SVN -> Perforce -> Git -> Git 3.0 -> [...] been-there-done-that. This is something that both open-source and commercial projects have had to deal with over the years.
Or maybe it would be a repeat of Python 2.x -> 3.x. ¯\_(ツ)_/¯ With AI assistance, hopefully porting the tooling over may go a lot quicker and smoother.
[0] https://www.youtube.com/watch?v=eJJp0RE7cd4&t=1134s
Her entire section 2 (starting at 6:12) is basically about why git does not and will never allow mixing of SHA1 and SHA256.
The interop discussed is using copybara as a copy tool to move data from SHA1 based repos to SHA256 based repos and vice versa.
My point is not that it's the end of the world (or the end of Git), but that it will be painful and unclear and confusing to lots of people. That would be fine if it made a huge difference in trust or protection, but it's the wrong way to do that.
Hm but the date is stored inside of the commit. The only way we can know that a commit's date is authentic... is through its hash. If I can forge commits with any SHA1 hash at will, I can make a repository whose head commit has the same SHA1 as the one in torvalds: /linux but where any commit was replaced by a malicious commit with the same SHA1 and a fake date. You have no way to detect that my repo is inauthentic other than through a deep history comparison. The whole idea behind a merkle tree is that just checking the hash of the top is sufficient to know the identity of the whole tree.
I don't know what the solution is, but I'm inclined to believe that any repo with a single SHA1 commit is as weak as a repo with all SHA1 commits.
If there is one way enforcement (i.e. there is one point where the last SHA1 commit was signed by first SHA256 commit), I think it should be safe ?
The "commit before" might be compromised, but the git commits refer a snapshot of a tree + a list of previous commit IDs, so the "new" SHA256 commit will not have any files altered
Every file (indirectly) referred to by a SHA256 commit using a SHA1 hash in some tree object can still be spoofed. Fixing that requires rehashing all objects and recreating all tree objects so the tree objects referred to by SHA256 commits are purely made up of object references computed by SHA256.
The date that a repo receives a commit is known to that repo. And a repo can stop accepting new SHA1 objects. And a SHA256 object could have a flag that says that no SHA1 objects may ever reference it.
The design of Git, as a Merkle tree, is meant to allow for use cases like this:
* I host a mirror of the Linux git repo.
* You download Linux from my mirror.
* You check out a commit, say fd179f8a05be3ccae366b9b96e176b51fbe54aab, which you know is a genuine commit through some out-of-band mechanism (mailing list, GitHub web interface, a line in a Nix file, whatever).
* You check whether the repository I gave you is legitimate or not by re-computing the hash of the commit which I claimed was fd179f8a05be3ccae366b9b96e176b51fbe54aab. If it comes out to be fd179f8a05be3ccae366b9b96e176b51fbe54aab, you know it's legitimate. If it doesn't, you know it's fake.
This is a completely normal use of Git. People download from mirrors all the time. People rely on commit hashes to identify a specific source tree. People trust that if whatever the mirror gave them hashes to the right value, it's genuine. That way, you don't have to trust the mirror.
If I can forge my own commits to have any hash I want, this whole model breaks down. I can replace some old commit in the repo with my own forged commit with the same hash, and when you download a copy of the Linux repo from my mirror, you'll receive a repo with malicious content, but it'll hash to the same fd179f8a05be3ccae366b9b96e176b51fbe54aab hash as a genuine repo would. This breaks the security model of Git.
Do I understand correctly that GitHub currently doesn't support SHA-256 repos at all, so the author faked a screenshot of what the create repository UI would look like if they choose a really stupid way to add support for SHA-256 repos, and then proclaimed it an unsolveable problem? Selecting the repository format based on the first thing pushed to it is really not a crazy idea.
The submodule problem is real, but the forge problem is entirely just "if forges implement support in a way that makes it painful it will be painful", and that's true of literally every feature that requires forge support.
I did generate that screenshot as I am not in the beta for this - I don't think many outside of GitHub are. However, this is exactly how GitLab does it and I would be _really_ surprised if this is not almost exactly how it's implemented. There is almost no other reasonable way to do it.
You can't select the repository format based on the first thing pushed to it, because both sides have to be initiated before a transfer can happen.
So you either start the project on the server and then clone an almost empty repository to start working (which I think is rare) - in which case you need to choose the format like this.
Or you initialize it locally and start your project and then push it to GitHub at some point, in which case you need to initialize a server side version that matches the format to push it to. There is no "initialize a new thing on push" and there never has been.
I personally would have preferred if there is a proper disclaimer that the screenshot is an approximation of how the UI would be. Or alternatively, use a more distinct art style of conveying the UI (e.g. comical/sketch lines or whatnot).
Right now its only a vague statement of "will need to look something like this", which does not imply on the originality of the image. I myself am misled that the screenshot is a legit UI.
The caption literally says "will need to look something like this"
> both sides have to be initiated before a transfer can happen
What would be preventing a forge from answering with a reference to schroedingers octocat when inquired about the yet-undefined properties of a newly created repository? The first pushing client explicitly looks for it, and no final decision has be made server-side until someone wants to take a peek, no?
> There is no "initialize a new thing on push" and there never has been
Well, there hasn't ever been a real reason for it, and now there is?
I also don’t know how big this problem really is. You practically have two ways of creating a repo. Local first and push or create a repo remote with Readme etc and clone.
The submodule issue is a different beast. But that one is currently a problem as well when a person choose a http url or ssh url. I have some extra git configs to normalize everything to ssh for instance.
github does support sha-256 repos in some kind of private beta. https://github.com/bk2204/talk-rust-in-git is one of them.
but yes: that ui in the article is likely entirely imaginary, but it's likely an option (either in the ui for private beta users or by raising a support ticket or by some internal tool) to get sha256 repos today.
I completely agree with everything you say here. I felt the critique in the article was way more strawman falacy than everything else. I use(d) git much more locally, and would never run into this problem because I would only use sha-256 on new repos. Rewriting the entire history is also not so much of a pain depending on organisation size and how it is done (if an org decided to do it). It may be a pain for git-hosts, but even there i think reasonable solutions can be found. Still not sure if I care about the switch though sha-1 worked and the security implications seem minor to my use-cases. I am much more concerned with code leaking online...
sidenote: i hate the term "forge", is it too late to settle for something less... cringe??
I think this change is more to do with politics rather than "security". Those kind of things where companies or gov, need to be certified with those super secure certificates and can't be using software that uses SHA-1. I don't have proof, but I'm not doubting it either.
This is what I saw in one of the mails. > > There are organizations where SHA-1 is blanket banned across the board - regardless of its use
And also on git 3.0 breaking changes. > > SHA-1 ... recommended against in FIPS 140-2 and similar certifications
Since SHA-1 isn't used for security in git, they should've instead moved to a non-cryptographic hash function such as MurmurHash3 and avoid all these problems, instead of moving to SHA-256 until SHA-256 is broken and need to move to the next cryptographic hash that is now incompatible with previous versions of git repositories.
SHA-1 is used for security in git. It's the thing that guarantees a commit SHA is unique. Without that, you open up all sorts of downstream infrastructure to supply chain attacks, where old objects get replaced with malicious ones, and then replicated on each subsequent git pull.
Linus' old argument was that the substitution would probably be noticed eventually, but that's specific to the way Linux uses git, and what he said probably isn't true in practice -- even if it is, there have been enough supply chain attacks since then to prove that even temporarily serving the wrong stuff to developers or CI is enough to allow lateral movement into other packages, production machines, etc, etc..
LWN had a good write up on this a while back: https://lwn.net/Articles/715716/
> At that time, Torvalds responded that SHA-1 is not the real security mechanism used in Git and, as a result, even a full compromise of the hash function would not necessarily be a problem.
It seems that the usage of SHA-1 is interpreted as a security mechanism while Linus used it mainly for other reasons, such as look up speed and deduplication of objects.
You can read the original README file when Linus created git[1]:
>+TRUST: The notion of "trust" is really outside the scope of "git", but
>+it's worth noting a few things. First off, since everything is hashed
>+with SHA1, you _can_ trust that an object is intact and has not been
>+messed with by external sources. So the name of an object uniquely
>+identifies a known state - just not a state that you may want to trust.
> ...
> +Another way of saying the same thing: "git" itself only handles content
> +integrity, the trust has to come from outside.
Yes if SHA-1 is broken, then content integrity can be broken but to me it looks like Linus at the time looked it from the point of view of corruption of files instead of "malicious" files.
[1]: https://git.kernel.org/pub/scm/git/git.git/diff/README?id=e8...
> There are organizations where SHA-1 is blanket banned across the board
This is very likely the case. And if it is, then it's a lost battle. You simply can't reason with that kind of corporate people, let alone have an argument around this level of complexity. Kafka (the writer, not the message broker) predicted this 100 years ago.
When going through the article, my instinct was changing from "annoying" to "this really sounds like a Python 2/3 moment for Git" to finally "oof this is going to be a mess" in the libraries/submodules part.
Meanwhile they are looking at us and thinking "you simply can't reason with programmers, they insist on using broken encryption"
Haha that's funny even if it's not accurate.
In my experience corporate box checkers don't care about reasoning (much less "encryption") at all, they see it as an annoying blocker in their path to the next promotion.
From what I'd read, SHA256 in git is showing every sign of being another IPv6. In particular:
- It's implemented in a non-backwards-compatible way
- The benefits over the older model are a bit nebulous
- There's a large amount of tooling that needs to catch up, and little sign that there is movement there
The difference with IPv6 adoption is that the internet relies heavily on network effects: so long as some hosts only have an IPv4 address, you need an IPv4 address for full connectivity, but then if everyone has an IPv4 address anyway, there is no immediate need to migrate to IPv6.
(Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting and peer-to-peer networking and so on; we are not the average user.)
This effect doesn't exist for the Git migration. Each repo can be updated independently; it doesn't affect users of other repositories, and most likely, the majority of devs will work on some SHA-1 repos and some SHA-256 repos with no issue.
If anything, I would compare it with the Python 2 to Python 3 migration, which was also painful, but succeeded eventually (despite being much less necessary in the first place).
You are wrong with your own examples.
Github is THE main platform for git. If github doesn't upgrade (and their code has been shit and hard to fix/update before) then the shift will not happen. Because yes, you can upgrade your repo independently, but if there is nowhere to push, no one will do it.
IPv6 is (also because of github) a great example for this. You can easily have an IPv6 address next to your IPv4 address, but a lot of websites (e.g. github) don't have that. Why would a normal company use IPv6 if even the bastion of nerds doesn't use it?
And to Python 2's "eventual migration" I can unhappily tell you, that my company (recently) bought an actively developed tool, that still uses Python 2.
I don't think GitHub is quite as important as that, outside of some specific projects that made bad/lazy decisions (thinking Golang here). If they didn't support 256 I think a lot of enterprises and open source stuff would simply jump ship to GitLab or elsewhere. In the enterprise world it only takes one person to write a security document banning sha1 for this to happen. For that reason, GitHub will support sha256.
GitHub is one platform that can make a centralised decision. What it does, goes. If only IPv6 had such an entity, we'd have it by now.
> Yes us Hacker News users have plenty of use cases for IPv6, like self-hosting
Funnily enough, self hosting is why I can't use IPv6. I want vlan isolation, but only get a /64 from my ISP.
Fortunately the lack of IPv6 also isn't a meaningful loss anyway so whatever
What kind of isolation are you looking for? Vlans might still be possible.
FWIW there's nothing about the addresses themselves that stops you subdividing a /64, but it depends on your router.
You could put your server at x::1 and static-route that address as a /128 on your router, if it supports it. The reverse route might be a bit tricky but putting ::0 on the router and telling the server it's a /127 should work. Anything outside of the /127 (so, all the randomly generated addresses on your home network) would go back through the router.
Now if that /64 is also changing every day, then it's a problem and idk what you'd do.
I just started learning IPv6 with AWS since they charge $0.005/hr per IPv4. Maybe it will be more expensive in the future and eventually it will be the new default.
changing repos to the new IDs would break any existing links to content on the pre-migration repos.
Unless you link using tags.
It could also be another Python 3 situation: backwards incompatible, unclear benefits with many downsides (3.0 and 3.1 being very slow), large number of libraries that need to catch up.
Does anybody know why it wasn't implemented in a backwards compatible manner?
One could wrap a whole merkle-tree with an additional extension tree, that just adds the new hashes. That way both kinds hashes could be used to traverse all data. The new hashes could be used to check the consistency, the old hashes would still be there to use in UIs or old release documentation. The downside being that you introduce more nodes in the overall data-structure which will have to be supported basically forever. And if SHA-256 is to week a third layer would need to be introduced. But the point is, it could be done. Albeit it would loose some of the elegance of the data structures involved.
Allowing both hash algorithms is, from a security standpoint, equivalent to just using the less-secure hash algorithm.
All repos need to end up using SHA-2 exclusively by the end. All tools that speak only SHA-1 need to be made incompatible intentionally. If the SHA-1/SHA-2 hybrid approach could allow that to happen, then it would be useful. If not, then it would just be a waste of time.
The way you would do that generally is to make it backwards compatible, then adding warning to legacy tools, then turning SHA-1 off by default, then removing it entirely. Doing it in a backwards incompatible way creates a chicken and egg problem, can't convert repo to sha-256 because some tool doesn't support it, tools don't have an incentive to be updated because no repositories.
> - The benefits over the older model are a bit nebulous
This is far from the case with IPv6!
Is it? There are many benefits in principle to IPv6, but if my ISP continues to assign me a single dynamic IP, those benefits are entirely moot for me.
If your ISP is assigning you a single IPv6 address, they are doing IPv6 wrong. You should be getting your own /56.
Your ISP assigns you a single dynamic /128? Which ISP is that?
Is the argument that there are no benefits to IPv6 for anyone/society because your specific ISP messes it up?
Right, but just because you don't happen to have IPv6 right now, how does that remove the benefits for others to have IPv6? That's like saying having a faster CPU wouldn't mean faster performance, because I don't have that CPU yet.
Do they actually or do you just assume they do because your computer gets one address?
There's a massive push right now from top down to have secure software supply chains. Google SBOM and SigStore. It's not an organic need but if you have government customers you don't have many options.
Ironically rewriting git history is a perfect opportunity for a supply chain attack.
The switchover date is usually announced well in advance and any interested parties can easily verify that the conversion was authentic.
I thought this would be a snark but it's an extremely well put together argument against the "Hashmageddon".
If you're replacing the weakness of SHA-1 just by going to another algorithm, you better be prepared to go to the next one when sha256 collisions happen, and it doesn't sound like git's design would be easy to modify for this type of crypto agility.
I do like their proposal for using signatures to establish trust and allow swapping sha256 for whatever comes next.
Technically, git's design (thanks to very smart people trying to solve this problem like brian and others) is _very_ easy to modify to different hashing algorithms now. A lot of amazing work has gone into this in recent years.
However, it's not a git problem. It's an ecosystem problem. It's that every git repo has to choose one and they're entirely incompatible with each other. That is the cost and the difficulty.
Git should support multiple hashes for commits
I think there is a question though when that will happen and if it will be in our lifetime. SHA-1 started showing weakness in 2005 (collision in 2^69 instead of expected 2^80. This was later brought down to 2^61 in 2011), the same year git was invented. Nobody has found a similar weakness in SHA-256 as of yet. SHA-256 is still at its design strength of 2^128
It took 20 years to go from vulnerability in sha-1 to having to replace it out of caution. There is no such vuln in sha-256 yet. It could easily be 25 years before we find one, and another 25 years before we have to do something about it. Perhaps longer. Will git still be used 50 years from now?
With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.
All it takes is just one collision to consider it broken right?
But hey maybe the attempt to fix it makes git controversial enough it falls out of favor, and nobody uses it anymore in 2 years, problem solved? sure.
> All it takes is just one collision to consider it broken right?
No, its considered broken before that stage. i.e. when someone discovers an attack that would allow someone to create a collision faster than they should while still being impractical.
> With the kind of compute power available nowadays and AI models I wouldn't be surprised we see it much sooner.
Computer power doesn't super matter, what matters is algorithmic breakthroughs. So far i dont think there are any examples of major breakthroughs of that type via AI, although perhaps i am just misinformed. Its still early in the AI revolution, it might still happen, but as it stands i don't think there is any reason to worry about that.
Perhaps not clear enough, but Scott was a cofounder of GitHub, so he knows a thing or two about git in the real world =)
Being a founder of GitHub doesn't make my opinion more interesting. I hope the argument stands no matter who wrote it. :)
It does though, even if it shouldn't be blindly taken as gospel. Arguments help, but some folks have a better brand of apple box to stand on, and "I live and breathe git" helps quite a bit ;)
This means, if you migrate your repo, every single commit message that contains text like: "please see commit <sha1>" will now be broken.
This will be a train wreck. I hope they don't release before adding compatibility modes to keep the existing sha1's around in the database.
Tools like git-filter-repo[1] support rewriting commit hashes in commit messages. git-filter-repo actually does it by default; see `--preserve-commit-hashes` in the manual[2].
[1]: https://github.com/newren/git-filter-repo
[2]: https://htmlpreview.github.io/?https://github.com/newren/git...
Sure, but that's not going to rewrite Slack messages, emails, GitHub links, docs
Came to say something exactly like this... The old sha1 handle needs to be still available in the same way an HTTP 301 redirect would work.
I run into broken documentation links all the time at work. A few more isn't going to break us.
... do not migrate old repos? I'm not sure why people would do that. Or, if they do, why would they replace the current repo name instead of creating a different one and keeping the old one closed to make the references work.
I don't think this is going to be a problem at all.
I need git-filter-repo to rewrite entire documentation and also resurrect and rehire earlier employees to repeat their GPG signatures.
Note that there exists multiple ways to continue to lookup SHA1s in a SHA256 repo.
One example is to maintain git-replace refs for the rewritten SHAs but there also exist config flags to enable object format compatibility extensions that help translate the SHAs back and forth.
https://git-scm.com/docs/git-replace
Once this starts being actual pain, we will each vibe the replacement index creator (git already supports replacement objects), for back-forth conversion, populated on pack and object indexing.
For massive perf and mem use damage. But oh well. And then we will wait for official version
There are plans to keep sha1s around in a database, but as far as I know, no way to transmit those, so they seem specific to individual forges. They can be recomputed, sure, but again, any signatures break and it's possible that in the case of an actual replacement, the recomputation is now wrong and not easily comparable. So what is the point?
This is not about forges, it is about the repo I have in a folder on my computer. The sha1 hashes shouldn't go away. Yes the forges also need to support this.
Since SHA-1 is already broken (just expensive in terms of GPU-time), then the text "please see commit <sha1>" is also already broken.
You can't attack an existing normal commit.
But also collisions there aren't a big deal. People will cite short hashes when referring to things and that's not "broken".
If it's an existing commit and you're already converting the repo, you can just convert the commit messages as well.
That is essentially only a second preimage problem, which is basically impossible.
Tbh, I would have imagined the SHA-256 transition to work differently.
Let's treat git as a SHA1-keyed object storage. The problem is that we currently use SHA1 both as database key and as hash for integrity validation. At first, I would have only changed the latter.
Local: Request object with key x (SHA1). Remote: Here are the bytes for key x (SHA1) with hash SHA-256. Local: Validate the bytes vs SHA-256.
Local: Store the following bytes with key x (SHA1). Remote: Check if key x has ever been stored in the database. If so check that the SHA-256 of the new bytes matches the SHA-256 stored under key x. The only thing the repo has to keep is a LUT from stored keys (SHA1) to hash (SHA-256). The prevents SHA1 collisions from being stored.
The SHA1 key then just becomes a convenient alias for an object. The only restriction is that you can't have two objects with the same SHA1 in a repo. We already kinda do this when we refer to commits with the first few chars of the SHA1. If there is a "collision" git already detects it and asks you for more characters.
Finally, you can convert the internal representation of the tree to SHA-256. If some legacy client requests aliases via SHA1 you use the LUT, new clients request the SHA256 directly.
Am I missing anything obvious here?
I positively don't care about the collision issue.
If you need to certify the authenticity of some code, and you've decided that a Git hash of any kind is going to be your certificate, you have a problem between keyboard and chair which is not fixable by stronger hashes in Git.
I don't want instability and churn in tooling.
It was about 18 months back, with gitea, I was starting some new projects and went sha256 because I’m a nerd that adopts things early. I could do basic git stuff, but I was stunned by how much didn’t work. A lot of actions, maybe even the entire action runner itself wouldn’t/couldn’t/didn’t work. It’s a deep cutting change and an expensive one that doesn’t have a new feature value. I abandoned that effort and rebuilt my repos.
Intentional attacks and collisions aside, what do you say to the fact that Sha1 had a lifespan at design time? NIST has issued retreating usage guidelines for the last 15 years and they themselves say it should be completely phased out by 2030. It just seems like good hygiene to switch it out. I agree, it will have a long tail and be a big bowl of suck, especially for the not dead but rarely touched code.
I appreciate the intent of “modern” tools with no switches, no ways for devs to make bad security choices because the hash, cipher and encryption mode are fixed vs all the crazy looking asn.1 stuff in tls and openpgp to enable n-degrees of configuration but this very issue is the counter example.
> what do you say to the fact that Sha1 had a lifespan at design time?
Read the article. SHA1 was chosen for content hashing, not security. Hence the security lifetime doesn't really matter.
So they’ve been talking about this for many years, planning, and finally announce when they’re going to switch the default.
So this is the right time to post that everything they’re doing is wrong? Did you engage in all the discussions about it and how best to handle it? Whether SHA-256 was the best solution?
I don’t see anywhere that it talks about alternate proposals or why they might have been better. Why the particular suggestions here were rejected.
This seems like a bunch of Monday morning quarterbacking.
I do mention this in like the first paragraph. I don't feel great about it, but I've listened to these issues for years now during contributor summits and Git Merge talks and while it's always seemed problematic, I thought they would come up with a good solution. This last Git Merge confirmed that it's close to the switch and not in any way solved or improved. I don't want to just go with it for groupthink reasons. I never thought it was a good idea and I have said that, but we have a last chance to rethink this, so I'm curious if I'm alone or in the silent majority.
Your argument is persuasive and well illustrated. I think the problem is the intro paragraphs come off as too certain of catastrophe which, when juxtaposed with your claim that "smarter people than me have been working on this", makes it sound like you don't actually believe they're smarter than you. The rest of your essay feels fair and not judgmental.
I do believe they're smarter than me, but sometimes very smart groups talk themselves into ultimately impractical solutions because they're all smart. Sometimes you need a dumb guy to come in and say "are you sure this is right?"
The plans have been on display in a cellar. Beware of the leopard.
This isn't really applicable. The plans have been talked about for a while very publicly.
I don’t understand what this is supposed to mean.
It's a hitchhiker guide to the galaxy reference, where sure, something is technically available but not clearly published and there are hoops even for those who know what they're looking for.
(No clue if it's applicable here, I'm not aware of this case, but I believe that's the reference if it helps :)
Edit : exact quote, as Arthur's house is about to be demolished for a highway bypass:
"But the plans were on display…”
“On display? I eventually had to go down to the cellar to find them.”
“That’s the display department.”
“With a flashlight.”
“Ah, well, the lights had probably gone.”
“So had the stairs.”
“But look, you found the notice, didn’t you?”
“Yes,” said Arthur, “yes I did. It was on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard.
I believe it's a reference to the hitchhikers guide to the galaxy – where the plans to remove the protagonists building to build a bypass road was hidden in this way.
nofunsir is a stoichastic parrot, matching to a bit in Hitchhikers Guide to the Galaxy, wherin the protagonist should have known to protest a plan to demolish his home where plans where clearly documented in a hard to find place that they could not have known about. It is not a good pattern match, because git has been discussing this in public on documented mailing lists for years.
A number of years ago when I heard about this, I was pretty angry and made a private fork of git immediately in which I tried to scrub away the SHA-256 bullshit. But that's basically just paddling upstream with a spoon for a oar.
The stewards of Git are going to do whatever they want, and there is nothing you can do about it if you don't have the clout to create a fork that takes the lead.
No amount of discussion will do anything because they've already decided that their view of the situation is correct. Git hashes are not just content identification but a digital certificate mechanism, and their collision resistance is a grave issue that must be fixed, the end.
You will be browbeaten in any discussion; it's not worth the energy in a world replete with issues.
This is actually a good change. If you want to change the security assumptions of Github repository, ie make them somewhat distributed. Then the SHA-1 based commit hash is a major problem. It only costs about 10k in 2024 to find a collision to a random SHA-1 hash. While this costs essentially makes the attack infeasible for most threat models. It does limit how far you can scale this without an obvious footgun waiting for you.
This is a good change, even though there is a massive technical debt in changing such a widespread system. It is worth the effort. Should generations from now still be using SHA-1 for their Git ops? Sometime you have to do the switch, otherwise you will never progress.
For my usecase basically i needed to know that every Git commit pointed at a cannoical blob. With SHA-1 you could generate two blobs which hash to the same SHA-1 hash, while you can do the format verification which helps i could not do that in my usecase as i did not know the underlying data. To fix this i had to very ugly have two methods of referencing any Git object, a cryptographically secure SHA-256 ID and the Git ID SHA-1.
Yes every repo is either one or the other but you fix that by rehashing the entire repo. Everyone can do this independently. It's entirely possible to maintain to identical repos in SHA1 and SHA256 mode but for the most part I suspect once updated people will simply pull down the new repo and use git 3.0 as a required version.
As migrations go, it's reading as simple to me. You'll just have to backpoint the commit signatures. I must assume there's a backwards compatible reference for them in git 3, right?
Or drop them and reference the old structure in a dire pinch.
Re-hash the entire repo as in rewriting all history? Hooo boy will that be a mess, I deal with things which reverence commits by hash in repos all the damn time. There are thousands of them in every Yocto project!
Yes, this would cause big issues for Nix based build systems or any others that reference commits by hash.
Do you have any external references to any commits that matter, for example in your communication platforms (emails, Slack) or your bug tracker? Or, worse yet, in places where they aren't just text format references, but used for things like CI/CD caching decisions or security scans?
Once you rehash the entire repo, every single one of those external references will be broken. Because no, there's no support for looking up old hash -> new hash or the reverse.
I think I remember something for this for mercurial to git migrations, I hope when its git 2 to 3 something similar is made (or it will take some time to adapt like when python did its 2 to 3 migration).
Since you're rewriting history anyway, couldn't you add the sha1 hash to all the commits as metadata?
"The migration to the new format is simple; Just re-write everything in the new format, but also keep the old format around forever too since data is lost in the new format!"
>it will be an incomprehensibly expensive and ultimately valueless and avoidable global nightmare.
thought "costly" in the title and "incomprehensibly expensive" in the subheader meant this piece would discuss how much less performant sha-256 is on modern machines, but didn't see anything. isn't there hardware acceleration? how much worse is it?
Actually, I think sha-256 is possibly faster than the sha1dc variant that Git currently uses.
I just sent a patch series to the list that enables sha1dc to be accelerated on modern CPU architectures to close to normal SHA1 speeds, but since it was ported from a Rust project by an agent, it will never be applied.
https://lore.kernel.org/git/20260929112544.86511-1-scott@git...
Last I checked SHA-256 was faster than SHA-1, and SHA-512 was even faster (though the output is annoyingly long).
How could SHA-512 be faster? It does more rounds of the exact same operations as SHA-256 with a bigger state. Although if it really were faster, SHA-512/256 gives you a truncated version.
[flagged]
What part of the code did you find bad when you reviewed it?
He means costly in terms of human effort and wasted time.
So many comments and none yet mentioned Linus Torvalds calling the SHA-1 deprecation a pointless churn: https://www.youtube.com/watch?v=sCr_gb8rdEI?t=11m
I suspect everyone renaming their branch from master to main caused more unnecessary breakages and toil than this ever will.
The problem is existing repositories have SHA-1 commit hashes, and if you also end up changing those...
There's lots of tools that refer to git commit IDs. Some of those tools may even hardcode a commit ID to be 40 hex digits long. The fact that these tools are external also means that "oh, just rewrite the commit messages or code to refer to the new IDs" isn't feasible. The only way to not break the world is to let people refer to existing commits with their SHA-1 hashes in perpetuity, and it doesn't sound like git is set up to allow this in any way, which means that existing repositories have to stay SHA-1 in perpetuity and that will cause fun down the line if you start having to make SHA-1 and SHA-256 repositories.
Changing from master to main is a one-off change. It might require changing your scripts once to refer to 'origin/main' instead of 'origin/master', but other than that, there is essentially nothing more that needs to be done, there is no risk to historical artifacts that needs to be mitigated.
> Mathematically, for SHA-1’s 160-bit output, the birthday bound means that you would need about 1.4 septillion random files (1.4 quadrillion billion files - 1,400,000,000,000,000 billion - it's impossible to effectively describe) in a single project to have file hashes accidentally collide.
The birthday problem gives the number of hashes to have more than 50% probability of having a collision. To have a collision, only 2 hashes are needed, with very low probability.
That's why it says "about".
Regarding the submodule question: I decided to test it and came up with this script: https://gist.github.com/Meneth32/bdb17a8787fc0ab197805d8d0cf...
On debian 13 (trixie) with git version 2.47.3, it runs without error. You CAN use sha1 submodules in sha256 repos.
I don't agree that SHA-1 is much longer feasible for git.
But I also don't think that switching to SHA-256 must be painful. A git2->git3 converted repo could just store all the past hashes, so existing links don't break.
I was thinking the same. Can't the git CLI see a hash and say "well, I don't see any matching SHA-256 hash, but let me check the Legacy SHA1 hashes I have stored", and still resolve an old SHA1 hash to the correct commit?
It could even do that + barf if it saw two objects with matching SHA-1 but mismatched SHA-256!
It's Python 2/3 or Node CommonJS/ESM all over again. I thought the lesson of history was that it's more reasonable to live with a known problem (and have mitigations for it) than to "fix" it in a way that breaks everything that's good.
I'm honestly still shocked any of this happened.
Prior to SHA1 we had MD5, a decade earlier. MD5 collision attacks had already been widely documented and known. It was the most obvious thing on Earth that this would happen to SHA1 too. Apparently, Linus never realized there was a need for cryptographic security and that the hash was purely internal.
Here's what I honestly think was a factor. I think C programmers fell in love with the implementation that you could throw around a fixed hash record on the stack. It's incredibly efficient. But it's an efficiency that doesn't really matter because as soon as you read from or write to a disk or a network or even memory, any cost saving is completely gone.
More than a decade ago, some people wrote a Java implementation of git (jgit?) and despite all their optimizations, it was (IIRC) only half as fast as C git. It is of course because Java at the time had no concept of stack values for non-primitive types so couldn't compete. Personally, I was impressed: only half the speed? That's pretty good.
For something that's only 20 years old, the Git SHA1 assumption is some of the worst technical debt we have in the modern era.
Here's another thought: when people make a lot of these programs, they often make the mistake of not separating the program version and the network protocol (or just the external API). So you end up with brittle client-server implementations where you have to upgrade both the client and the server at the same time because they lack a network abstraction.
The other end of the spectrum is video streaming where you have codex, container formats, transport protocols and so on.
What a mess.
Having gone through multiple "security" upgrades over my decades-long tenure, I firmly agree with schacon here (this post is mostly for schacon, since I'm reading a lot of negative on this discussion (Rubyists need to stick together)).
The hashing algorithm's security properties are a security property of Git. This is because:
* we pin to commit hashes and expect this to refer to immutable content,
* commit and tag signatures are over the hash.
The Linux quote about trusting the distribution doesn't make sense to me, as Git is content-addressable and decentralised, although possibly at the time it was a reasonable position to take for kernel development, but [1] could equivalently happen for Git and the hash algorithm being non-broken is required for it to be noticed. Not having to trust the forge is a very desirable property.
[1]: https://lwn.net/Articles/57135/
My prediction: 20 years from now everyone will still use SHA-1 git. That will be simply easier.
GitHub could make it the default for new repositories created via web UI and devs who don't touch the CLI probably wouldn't even notice
Once people realize the fundamental incompatibility of these new repos with the old ones, they’ll be furious. GitHub will back down.
It’s like IPv6. Just worse.
Do you want to bet on that prediction
I would
I think this is really selling the severity of "SHA-1 is a Shambles" short. The author distinguishes between second preimage and collisions, but Shambles is a chosen-prefix collision, so it's sort of in between the two ends of the spectrum he explains.
The difference here is between me being able to send you a benign file and then swap out a safe file (the scenario the author describes) and me being able to swap out your own familiar, benign file except this time it has malicious content at the end (along with a bunch of garbage).
That second scenario is obviously much more dangerous from the perspective of human review and noticing that something is wrong. If you have a large, rarely changing file already, then content silently smuggled onto the end of it could stay undetected for a long time.
(In practice, it's not that simple; you'd have to actually figure out how to make multiple pieces line up with chosen prefix collisions, and I haven't worked through how you would do it. Maybe it's impossible without further weaknesses. But that guarantee is feeling pretty threadbare.)
Anyway, this doesn't matter because git switched to sha-1dc back in 2017, and afaik, this addresses the shattered/shambles problems. I find it odd that this isn't the point the author focuses on; instead, the only mention of sha-1dc is in a footnote saying that it could be replaced with his proposal.
And sure, maybe that's true, but why isn't that the entire argument of the article then?
Well, this could change at any time. Part of the point was to argue as if SHA1 was completely broken and I still feel that its the correct argument.
But from your chosen-prefix argument, having some object thats been around a while is a problem, because any viable attack needs to be fairly new, since git wont replace objects it already thinks it has. Maybe fresh-shallow-clone scenarios like GitHub actions, but that still always has a non-sha based authentication protection (in other words, actions never run on untrusted code and that trust is never based on signed artifacts but on source provenance)
akschually! I argue in a footnote at the end that sha1dc is an unnecessarily expensive shim protecting codebases in a similarly unnecessary way and should be removed so that our pushes and fetches can be much faster.
With the increased computation given to LLMs for the new development practices, wouldn’t any extra sha be a much smaller accommodation (then claude, chatgpt, etc…)?
The post's argument that hash collisions are irrelevant in practice is not convincing at all. Basically they amount to:
1. Collisions aren't as bad as preimage attacks
2. Even if you made a file-with-malicious-hash, how would you get people to pull it?
3. Other attacks are a bigger problem (social engineering)
(2) is laughable in a world with github. It's common for unknown people to submit pull requests to code bases, and for those changes to be reviewed and merged. For example, as part of reviewing pull requests, I have `git fetch`'d proposed changes to my local machine to check behavior on some additional test cases. "If you fetch it you're fucked" is unacceptable as a security boundary.
(1) and (3) are just tu-quoque arguments about other attacks being worse. The relevant question isn't how bad other attacks are, it's how bad this attack is.
The fundamental problem with collisions is that software often assumes they can't happen (or is not tested against them). Thus collisions can trigger bugs, or otherwise cause surprising behavior. For example, webkit figured the colliding PDFs demonstrating a sha1 collision would be excellent for unit tests, so they merged the PDFs into their SVN repo... which completely fucked it [1]. I don't know the exact internals of git so I can't comment on how you would get surprising things to happen, but "oops the file you merged was different than the file you reviewed" and "oops the repository got corrupted" seem entirely plausible.
[1]: https://www.reddit.com/r/programming/comments/5vyhy2/webkit_...
(2 counter) is impractical because all nodes of git will not replace objects if it thinks it already has it. So any attack has to assume this is the first time the node fetched, which is difficult before trust is established, which is difficult. This is part of the argument Linus originally outlined for this vector, which is that it only works for _very recent_ objects.
(1/3 counter) is not what I argued. I argued from the worst-case position that collision and preimages were theoretically cheap and fast. Even in that case, I feel my arguments hold.
The main issue here is that you assume you can replace an existing object with a replaced one, which you cannot. Not only that, but in all known cases, the sha1dc variant of SHA1 that Git uses will even _tell_ you that someone tried to do this, which singles out the source quickly.
couldn't github reject a push that contains an existing hash in the repo?
It doesn't reject, but it will not replace. Same for a fetch/pull. That is another issue with this attack vector (that Linus also mentions) - it has to be the _first_ time that a node has seen this object. It makes the attack even more difficult than it already is (in like 4 different major ways)
In the middle of vacuous complaints about slightly late tools and debatable threat models, I see one legitimate-looking concern in this article: submodules require either hashing algorithm and therefore force inconvenient upgrades.
Is it true? Where exactly the parent repository references hashes from the submodule repositories, and how could these links be generalized for compatibility?
Ugh, I didn't know that SHA-1 submodules wouldn't be supported in SHA-256 repos. That changes the transition from painless to a major dumpster fire. Having to maintain converted forks, and use different hashes from upstream is going to be a mess.
git submodules has always been a major dumpster fire. You’re honestly better avoiding regardless of SHA-256 incompatibilities
The majority of C++ projects I work with use them. Telling people not to use a core git feature isn't viable
The only reason C++ projects depend upon submodules is because C++ is just as archaic as COBOL these days.
I used to be one of the more vocal defenders of C++ but honestly, if the reason you’re defending git submodules is because there’s no better option in C++; then you’ve basically already lost the argument.
Nearly every other language has found a better alternative. Even Go has, and that’s largely mocked for its prehistoric approach to modern programming language design.
I wouldn't call it a core feature -- more like a convenience hack.
This article says if you manually enable the experimental new code today then “you can't push code to GitHub” and that old and new repos “cannot be mixed”.
The git transition document talks about keeping a bidirectional mapping between old and new hashes, and converting between the two when pushing to legacy repos:
https://git-scm.com/docs/hash-function-transition/2.55.0
I suspect that’s what gitbutler.com means when they say that, although broken today, everything “will almost certainly be fixed” when git 3.0 ships.
While I’m sure they make a good argument in the hash theory section, I’m less inclined to believe the “train wreck” part of their post.
If you upgrade your house and leave the roof off then yes, it will be a “costly mistake” the next time it rains, but if your plans say you intend to put the roof back on then I’m not sure how I am supposed to interpret a blog post warning of the perils of a roofless house.
Oh, agreed this sounds like a terrible migration path and shouldn't really be needed in the first place.
What I'm missing in the article is whether any Git server accepts replacing a SHA-1 identified object it already has. If it doesn't, then the distribution trust discussed holds, and keeping SHA-1 seems fine. Adding additional signatures seems fine for those who need transitive trust.
> The first thing that you'll notice (other than the much longer hash value) is that you can't push this code to GitHub, though that will almost certainly be fixed by the time Git 3.0 is released. In fact, that's probably the main thing currently delaying 3.0 entirely.
> But when you do want to push it to GitHub (or any host), you will need to tell them when creating the repository on the server that this is a sha256 project. Every project will now be in one bucket or the other and they cannot be mixed.
> This is immediately going to frustrate people because now they need to know what version of git they ran git init with and make sure when they go to GitHub to create the server repository, they choose the right one.
Github can fix this problem, and we have had siilar stuff in the past, when we had to switch repos. This is simple and doable, just need gradual work. I don't see the problem.
Well, Fossil has already moved from SHA1 to SHA3-256 and I haven't encountered any issues.
https://fossil-scm.org/home/doc/tip/www/hashpolicy.wiki
I find it disappointing how few people are addressing the proposed "Independent Tree Hash Headers" solution. It's probably the most interesting part of the article, but it's getting the least attention.
I came in expecting to disagree strongly with the article, but ended up agreeing more than I didn't (though I still don't 100% agree, as collisions are still an issue for mirrors). I find the concept of multiple hashes per commit quite interesting. It would allow mixing hashes in one repo, wouldn't break submodules, and tooling could be used to reject commits without any secure hashes for a gradual transition (like enforcing signed commits/tags).
Make git init use SHA-256 if git is invoked as git3, SHA1 if invoked as git2.
Plain git init could fail with a diagnostic: informing to use one of the two aliases or an option.
What people don't want is making git repos SHA-256 by accident and finding out later that they made repos not compatible with older git.
just make SHA1 the default in all cases unless the user specifies otherwise
I agree, but you're not going to sell that argument to a herd which has decided that git hashes are digital certificates which must be replaced with SHA-256, or the sky will fall.
Can someone more cyber-pilled than me explain what the actual risk with Git hashes being susceptible to collision attacks is? Obviously accidental collisions are problematic, but to my understanding the probability of that is still approximately zero.
Best I can tell, all a forced collision would do is let someone who already has control of a repo modify the history in a far from plausibly deniable way. Which in practical terms, they already could do simply by replacing the whole thing, because who's out here using git hashes as a security tool? Every pinning I've ever seen has been to tags (which can be modified at will), or hashes of the actual payload (which doesn't need to be the same as what git uses).
> who's out here using git hashes as a security tool
Among others, dependency management in Rust [1] and Python [2] sometimes uses references that work similar to https://github.com/rust-lang/rust/commit/ec999ed [3] to suggest one particular version of the project, authored by the specified maintainer.
Unfortunately, it means neither, unless you pushed it. The hash points to whatever the first person uploading it to github submitted. And the author/org name in the URL is window dressing: all the objects go in one big bucket regardless of push permission to one particular fork (because why wouldn't they - today, collisions are believed to be recognizable because the cheapest way to craft them results in clear tells).
[1]: https://doc.rust-lang.org/cargo/reference/specifying-depende...
[2]: https://pip.pypa.io/en/stable/topics/vcs-support/#git
[3]: N.B. the "This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository." warning Github has started to add to URLs like that.
I see commit hashes used all the time, like in yocto recipes for example
Why are we quoting a technical opinion made in 2005 as if 21 years of time passing has somehow solidified the opinion? I mean, maybe it has, but that's not how technical opinions work
I agree with the main thrust of the article, that we should instead trust the transport mechanism rather than the cryptographic properties of SHA1, but because of this I don’t really think switching the default will be a costly mistake. I rarely use full SHA1 hashes right now; I only use the truncated version and I don’t think any user cares about the length of the full hash. As for compatibility with forges, it’s just a small UX problem that should be solvable: don’t let the user choose the format when a repo is created; instead choose it when the first push happens.
I don't why a sha256 project could not support sha1 usages too: if you try to merge/rebase/cherrypick/whatever a sha1 commit in a sha256 project git could simply recompute the needed s256 data and keep old sha1 ids as validated metadata. you could even use either hashes as tree-like refs.
you could even do the same in reverse and "import" sha256 commits in sha1 projects. the only real difference would be which one is used as primary key in git's internal store.
Because it's slower. You would have to build an internal storage format and a network format that allows multiple hashes.
One of my theories (see my other comment) is that Linus and the people who took over Git fell in love with the implementation. A fixed size (160 bit) hash on the stack is incredibly efficient. A variable length record with optional fields isn't. You can do a byte-wise comparison of 20 bytes on a SHA1 hash. You can't if there are unknown hashes in there that might not break equality.
You should never let these kinds of implementation details drive the design at the expense of future compatibility. We were already aware of this issue in 2006. It happened with MD5 (to SHA1). So now we have some of the worst technical debt to be introduced in the 21st century and it's going to be a giant pain to fix. SHA256 will eventually fall by the wayside too so just changing the algorithm isn't the long term solution.
> We can go through years of this SHA-1 to SHA-256 migration and then quantum computers break 256 and we're back in the same stupid boat again.
SHA-256 is considered quantum safe by the NIST and is left out of PQC migration guidance entirely.
It was theoretical - the point was that maybe some paper is published or some new tech or issue comes up. Now we have to do this again. If we separate the concerns, then we don't have to deal with both as though they're one problem. We can deal with one thing for content addressing and another for trust and security.
Heard. If we’re going through this pain of a migration- adopting something that is forward compatible at the ecosystem/community level like ipfs makes sense to me. You throw a byte at the front of the hash which indicates the format the hash is in.
We have a content addressable storage system internally and we use a self describing hash for it, so we can change the hashing scheme for different use cases.
Multihash is ipfs’ container for this.
Until it isn't.
Things like this have a tendency to to be revised as time passes.
For now. Turns out the pigeon hole principle still holds.
1x10^77 is a lot of pigeon holes.
Yep. So was 1.46 x 10^48
A lot of replies here seem to be asserting that this "isn't that hard" without addressing the thing that makes it most hard: submodule compatibility and the breadth of tooling
Submodules are a mistake.
They're the best way we have to reference other repositories from one repository. All other solutions don't have the benefit of being built in to git and having support built in to all git forges.
Ecosystems like Yocto are built around having meta layers as submodules. And, despite the usability flaws of submodules, it works really well.
I also use submodules to include dependencies into C++ projects a lot. It works fine.
git subtree and git subrepo are compatible with all git forges and don't require normal developers to install the extensions. Only the person/bot doing the occasional sync to the external repo has to install the extension. I prefer git subrepo for most (but not all) use cases.
I won't defend submodules, but I also don't accept this as a response because it's irrelevant. They are used and it will be an unbearable pain when they break.
I find them extremely useful. It gives me a monorepo experience in repos that otherwise aren't/can't exist as one monorepo for various reasons.
> Submodules are a mistake.
Someone started this FUD a long time ago and it has worked. Instead of using an elegant mechanism, project have built inelegant wrappers on top of git like go.mod which are actual mistakes.
It is the same people who say Git is unseasonably hard to learn. It isn't. Git and its command line is ugly and hard to remember. Understanding what those commands do at a commit and a branch level isn't really difficult. Even the internal object model is only moderately difficult.
I say this as a person who strongly dislikes many many aspects of Unix and Linux due to bad design and terrible UX. Git has a better design than any Unix program you get.
Submodules work okay. It is just Git LFS but for Git repos. Get over it.
I’m not the author but I agree with their opinion.
Compare the UX of go mod with git submodules. One is easy and the other is about as fun as having teeth extracted.
git’s UX has never been its strong point. But submodules takes that pain to a whole new level.
Core of the issue seems like github UX issues that they can solve
In fairness, there's a second issue: Git repos should just support multiple hashes, so you could just add SHA-256 to a SHA-1 repo, which would then transparently support both SHAs.
This even would make the SHA-1 git objects collision resistant when stored on a trusted server, even with untrusted clients. (Exercise left to the reader.)
A question from a non dev and non sec.
As I understand, SHA1 hashes in first git gave us both: - 1. unique identifier, - 2. weak attestation that the commit body is not modified.
Later SHA1 was found to be weak. We still have 1., but 2. is not assured anymore.
Q: Why not give the paranoid people¹ their SHA256? In 2026, Emacs starts as fast as Vim. Is there a difference or difficulty in implementing SHA256?
AD 1 - discussion whether paranoia is justified is not important if it's a simple fix.
TFA argues exactly this, that SHA-256 should not have replaced SHA-1, but it should have been added as a second hash.
Nonetheless, this solution has its own disadvantages, which were considered as more important by those who have chosen the replacement solution.
In my opinion, SHA-256 should be used for any new commits, and all old commits should be rehashed, but the old SHA-1 hashes should have been preserved and stored in some format that would have allowed the retrieval of the new SHA-256 hash corresponding to an old SHA-1 hash.
Seems like rage bait, makes a claim of a lot of computation in the era of LLM where computation is already out the window.
I know this is just one tiny piece of the omnishambles, but: can't github sidestep the "you have to tell us what kind of hashes you have at repo creation time" by effectively deferring "git init" in practice until it sees what's being pushed? There may well be a good reason this wouldn't work...
It would make git compatible with newer projects that store SHA-256 hashes. For example Safecloud can be used for encrypted storage and streaming: https://safebots.github.io/Safecloud/
I feel like if it takes this much ink to explain why using an insecure primitive is actually safe, you should just fix it.
The arguments make sense, but how ironclad are they? How confident are you that some clever black hat won’t figure out a way to take advantage of it?
This is one of the most widely used programs in the world. Let’s close the hole.
Ok, so the guy manages GitHub and doesn't want to pay the price.
Fair enough.
But the worrying part is that he seems really convinced that his security analysis of the vulnerability is complete, and that we should be convinced as well.
If you didn't need a strong hash, use a very fast one then, like MurMurHash.
It seems people are missing the point: it's not even the submodule incompatibility that's going to become an issue majorly (like python2 --> python3 but worse), the main issue is the loss of traceability for repos that changes in place (which I assume many will do). Imagine what will happen to these:
- SLSA and Provenance or SBOM data in the supply chain security that uses commit hash. All the previous images are now pointing to a non-existing commit
- All the documentation and tooling as the article calls out
- All your traceability links from your project tool to your git repo, they will lose all the past data as it will be dead links
So I hope there IS NOT a migration path for in-place replacement!
The xz backdoor is basically his point in practice — that was a maintainer-trust compromise, not a hash collision.
You'd think after all the failed (i.e. IPv6), and barely-succeeded (i.e. Python3) migrations the open-source world has seen, that the architects thereof would prioritise backwards-compatibility a little more highly...
schacon: Really like the "Independent Tree Hash Headers" idea. How difficult would this be to get this functionality into git? Would it cause any breaking changes with older versions? Have you discussed this with any git devs to see if they are open to adding it?
Actually, this entire blog post came out of a short chat at Git Merge a few weeks ago with Jeff King. I argued more or less this and he didn't _entirely_ disagree, though he has good counterarguments on the list over the last few years, so I don't really know how he thinks about it ultimately.
I would write this to the mailing list, but I thought a conversation that includes people outside that list is more interesting to me. Ultimately I'm not sure if I'm dumb about this or the whistle blower that's willing to actually say "maybe this isn't the right call"
Also, functionally, this is incredibly easy to add to Git.
Meta: This site breaks the back gesture (2 finger swipe to left). Very annoying
Smells like "sky is falling" kind of breathless claims.
I'm no cryptanalyst, but I'm also not seeing this anywhere else. If it were such a problem, I would be seeing complaints elsewhere.
Git might be the technology with the lowest understanding per user score. Which makes it not surprising how major release of git flies under the radar of tech media coverage.
You're right on the money with this one. I'm one of the people in that group.
Have any recommendations?
So what about my self-hosted repos in gitea that use SHA-1? Will they all become unusable once I can't avoid updating to git-3 anymore?
I run a small git/jj forge and for us it's already painful dealing with this. Can't imagine how GitHub is going to handle it.
The first reason the author lists for why this will be bad is only an "issue" on Git hosts that don't allow repo creation on push (which is brain dead of GitHub). Any other host, you push your new repo, and it will see the hashing algorithm, and receive the contents accordingly.
Submodules is a legitimate argument against this, though I don't know how widely this feature is actually used, and similar to the arguments in favor of switching the default branch from master to main, this is simply a setting which can be changed.
I do like the idea of commits having both hashes, and am surprised that idea has not been explored further.
Generally though, I think the author's strongest argument is simply that the change isn't strictly "needed", and all the other issues presented aren't the strongest arguments against change.
There are many disappointing errors in this article that one would expect an author with such prestige to know better.
Friends don't let friends build decentralized content addressible storage on insecure hash functions. Nor do friends try to convince you to use insecure hash functions. Nor do they appeal to authority. They also warn you against signing the output of insecure hashes. Finally they don't bother making argument that it's not worth upgrading today because we might have to upgrade tomorrow too.
I see someone has raised 17 million for a new ai git frontend.
oof. I can only hope that as many as possible git clients/wrappers that people actually use stick with SHA-1. Solving issues related to that not having happened sounds like the least meaningful work one could be doing...
I'll probably switch to git-evtag then, and keep the old SHA-1 then. Same as Google.
Meh. Git 3 can still init with the old sha1 hash.
Also GitHub should auto-detect the format on the very first “git push” and handle it then. If they actually require config to be set appropriately during the initial “new repository” dialog that’s just bad ux.
As for this being a looming disaster for industry… it’s inconvenient. And lots of things will break. But we will survive and come out the other side I’m certain. We converted from the Julian calendar to the Gregorian calendar 500 years ago. Surely we can handle this.
> We converted from the Julian calendar to the Gregorian calendar 500 years ago. Surely we can handle this.
To be fair that was more of a time sync problem.
If you think of the Pope as a stratum 1 NTP clock, with bishops in each diocese acting as stratum 2 and parish priests as stratum 3.
The break between Julian and Gregorian was sort of like a UTC leap second.
> Also GitHub should auto-detect the format on the very first “git push” and handle it then
As far as I'm aware, that's not possible because of how the git protocol is designed.
GitHub could have both styles pre-initialized behind the scenes, multiplex the push to both, and keep the one that doesn’t fail.
In software everything is possible.
I'm going to completely ignore the first part of this post because I'm not interested in arguing about how severe the issues with SHA-1 are. I think it's accepted that there are flaws.
So given that, I'm more interested in the arguments for why migrating to SHA-256 is problematic.
The biggest issue I see, after skimming over it, is the submodule breakage for new projects trying to link to old projects. This seems solvable frankly, but is the only serious issue I see. Everything else will be worked out as software is updated IMO.
the sha1→sha256 flip is gonna break every script that assumes 40-char hashes. my own deploy glue does that lol. not looking forward to the grep day
I think the only good argument here is that sha is maybe not a security control for git. I think every other argument in this article is incorrect
a) it's relatively fast and impossible in a practical sense for two different files to accidentally hash to the same value.
That is silly. We are not worried about accidentally triggering. We are worried about intentional triggers.
I dont know why people always bring this up for hashing. In any other context it would be considered silly. If someone said, the chance of triggering a buffer overflow by accident is low, we would call that silly as we aren't worried about accidental triggers.
b) second pre-image vs collision. In a world of open source where we accept commits from randoms on the internet, i think collisions are just as relavent as second pre-image.
a) it's silly to stop reading at that quote because the following text deals with "identification"
b) so, you agree with the blog? "So, any realistic interesting attack vector therefore relies on a collision attack,"
> b) so, you agree with the blog? "So, any realistic interesting attack vector therefore relies on a collision attack,"
My reading of the blog is that they are dismissive of collision attacks. In context of git, i disagree. I think there are plausible attack scenarios involving collisions, or at least, just as plausible as second pre-image.
If you mean do i agree with the blog that impossible attacks aren't possible? well yes obviously, but i think that goes without saying.
they could just have taken the sha1 of the sha256.
Then `git fsck` wouldn't be able to tell if an object uses the sha1 scheme or the sha1∘sha256 scheme, so it may need to compute the both hashes to check an object's integrity. Also, an object name alone wouldn't indicate that the weak scheme shouldn't be used to check integrity, so a malicious sha1 object could be swapped in in place of an sha1∘sha256 object (if second-preimage is found).
So many people coping.
No it won't.
> In other words, hash collision attacks are maybe the dumbest possible way to get untrusted code on a system when unpaid open source maintainers and low-trust package management forges exist.
Yes. And we literally see this every other month with NPM
Thanks for writing this. I'd only been loosely following it and I hadn't realised how bad this is going to be. I have repos with tens of submodules and it's going to be a nightmare if any of them switch to sha256 in place. Not to mention I won't be able to use any new projects unless I rebuild my repo and all the submodules therein.
I thought the master to main thing was bad enough but this is going to suck. And just like the master rename it achieves basically nothing.
What is it about these projects that attracts people who just want to change things for the sake of it? Real engineering means coming up with a solution for backwards compatibility. This is just irresponsible and, frankly, a fuck you to everyone who will be affected by this.
« people who just want to change things for the sake of it » — Not for the sake of it, but to “make the world a better place.” The intention is noble (well, mostly, at least let's assume it is). Of course, there is disregard for history and her deplorables (or not so deplorables), comes with being “progressive”. Which is why I like Windows better (not 11 though, will probably have to go back to Linux at some point).
« What is it about these projects » — Maybe that they're “at the forefront.”
> I pull it from there because I trust that GitHub has its authentication game together enough that it's unlikely that anyone malicious pushed something there without the maintainer's knowledge.
Hahahahaha, that's some level of delusion
This seems like Y2K fud.
The alternative to making sha256 the default is to leave sha1 the default. Nobody changes to sha256. sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue, but github never implemented sha256 because they didn't have to. This would be a major problem.
This is very very easy to fix if you run into it.
1. Adopt git 3.0 if you can with sha256.
2. If you can't use sha256, set the config to put things back to sha1. Wherever you need to do this you probably already set dozens of ENV vars or settings, just add a new one.
Or write a 15 page analysis about how the above is so hard people will probably just find it catastrophic to even think about.
> sha1 is broken in 10 years. Suddenly everyone has to switch all at once on the same day because it is a critical security issue
If you read the OP article, the entire point he's making is that this would never happen, because a hash algorithm being "broken" doesn't matter in practice, because true supply chain security has nothing to do with file hashes.
It's easy for _one person_ to fix. It's not easy for the entire git ecosystem as a whole. GitHub, large internal corporate git repos, CI/CD systems, projects with submodules, etc. The second half the article explains all of this.
It’s already broken, but even though it’s broken it’s hard to generate git collisions because of the repo metadata. It’s easy to generate (for instance) standalone PDFs with identical hashes, but doing this with git in a useful way is much harder.
That said, it’s still a good idea to migrate to a more robust hashing algorithm. Defense in depth, etc. Just because it’s a difficult migration doesn’t mean it shouldn’t be done.
You can certainly do this, as I said, this is Google's backup plan. But defaults matter. People will start running this and getting repos that are uselessly incompatible with other repos, tools, libraries and server instances. Having it as an option is one thing. Making it a default will cause a lot of pain for people who don't want to care about this.
"defaults matter" is an argument for this change, not against it.
No, my argument is that the change should not happen at all and nobody wants it and it gains the community very, very little but the default change is forcing it on everyone and most will be _entirely_ unaware - now having to solve problems that are difficult to understand. Defaults also matter when they are the wrong defaults.
> Adopt git 3.0 if you can with sha256.
Who is "you" in the context of a distributed version control system? I think this is not just the plural you, but the unbounded you -- it's all people who not just interact with your project now, but who you hope may interact with it in the future. The question is what the cost is of committing a near-infinite population to this migration, not the cost of doing a single `brew update` on your personal machine, no?
Just clone the repo again. Jesus Christ, you act like the simplest thing in the world is some kind of insurmountable challenge.
For the record, Y2K was not fud. It was very real, in a long list of datetime problems that are to come. Further datetime problems are coming at scheduled dates.
It was definitely FUD. There was a real problem (date counters would roll over), but the impacts of it were so ridiculously overstated that it eclipsed any sane discussion of the issue. We had people at the time predicting that planes would literally fall out of the sky when Y2k hit, which was never a realistic possibility.
Considering one of my team literally was carrying a Motorola satellite phone on the NYE/D of 2000-01-01 in case our mobile or land comms went down because the public transit was running all night, we took it very bloody seriously.
If there was an interruption to services on that night in particular, that's a serious public safety issue.
We had tested our system, and integration tested with all the other systems we connected to directly, but any FMECA analysis would show you that there were failure modes that we couldn't mitigate.
So people on planes falling out of the sky? Probably no.
People being crushed in a railway station on NYE? Possible yes.
> which was never a realistic possibility.
Because a lot of work was done to prepare and fix potential issues.
That's the problem with deniers. When responsible persons take preemptive action to prevent tragedy, like with Y2K, the diners say it was FUD. When people don't take action, like with climate change, they say it wasn't important considering it's not them who's dead, totally discounting those who have suffered or died as a consequence. In summary, the deniers are so incompetent that they can't be trusted to correctly maintain a car, let alone civilization, considering they would never even the replace the necessary parts at the right schedules in their car.
> not really practical to exploit in any demonstrated way
like gitc0ffee?
[dead]
[dead]
…
What does OpenAI have anything to do with this?
Would the author feel the same if git had used MD5 instead of SHA-1?
I do actually literally write in this that if it was MD5 it also would not be a problem.
Fair cop.
The hash isn't the security, the distribution is.
https://lore.kernel.org/git/Pine.LNX.4.58.0504291221250.1890...
As linked by another commenter in this thread, Linus worked out years ago that even if someone inserted a malicious object into the kernel repo, it would at best be a nuisance and not a major concern.
They address this very theoretical. In short: Yes. Which makes sense if you don't treat the hash as a form of security against malice, especially in the case of attacks that are already impractical, which is the entire thrust of the article.
Follow the ethos of the quote master!
> We could be using MD5 and it would honestly probably be just fine.
Didn't GitHub broke the entire CI system by switching to "main" branch :thinking_face:
Anyway, I think it'll be the same. Tools will support SHA-256 quickly and we might have a migration program that converts SHA-1 repo to SHA-256 repo.
The only problem is that git (and related tools) will get twice as big...
Claude, make the hash use SHA-256 rather than SHA-1. No errors plsss.
Besides the possible implementation/deployment issues they will or will not face with this update, I can empathize with the idea that of not wanting to have a possible vector of attack in your system. Particularly today with AI being able to find novel exploits, I could see a future where a vulnerable hashing system leads to a malicious injection attack.
The author argues that "If I wanted to get untrusted code into Android, it's so much simpler to bribe or convince the maintainer of a popular downstream project" which is a really a red herring in this matter since that is literally a completely different issue that obviously no software update could ever fix.
Nonetheless I do agree with him in regards of how complicated and messy this whole process will be. Crypto migrations have been historically difficult, expensive and overall ugly, but not impossible...
https://nvlpubs.nist.gov/nistpubs/gcr/2018/NIST.GCR.18-017.p... page 58 for instance.