Packages: Maven, NuGet and RubyGems registries for every workspace
Three more registries in services/packages, each in its tool's own protocol, with g1t tokens, the same access, storage, free limits, billing measure, events, audit entries, read rate limit and repository linking as npm and Cargo: - Maven at /-/maven/<workspace>/: the standard layout, mvn deploy and Gradle publish over PUT with Basic auth, maven-metadata.xml made per artifact and per SNAPSHOT, checksums worked out on upload (a checksums table) and served beside every file. A release's files are written once; the artifact's metadata upload, a deploy's last step, publishes. - NuGet at /-/nuget/<workspace>/v3/index.json: flat container, registration and search, dotnet nuget push with the token as the API key, the .nuspec read from the .nupkg, delete unlists. - RubyGems at /-/rubygems/<workspace>/: gem push and gem yank, and the compact index Bundler reads (versions, info, names), from the gem's metadata.gz. Small readers for zip, gzip and tar (archive.rs), XML and the YAML RubyGems writes. The Ecosystem enum gains maven, nuget and rubygems.
| 1013 | 1013 | "g1t-kit", | |
| 1014 | 1014 | "hex", | |
| 1015 | 1015 | "hmac 0.12.1", | |
| 1016 | + | "md5", | |
| 1016 | 1017 | "miniz_oxide", | |
| 1017 | 1018 | "serde", | |
| 1018 | 1019 | "serde_json", |
| 1 | 1 | //! The packages service: the registries a workspace publishes to and | |
| 2 | 2 | //! installs from, beside its code (docs/PACKAGES.md). Container images | |
| 3 | 3 | //! first, spoken over the OCI Distribution protocol on `g1t.sh/v2/`; npm, | |
| 4 | − | //! Composer, Cargo and Go after. | |
| 4 | + | //! Composer, Cargo, Go, Maven, NuGet and RubyGems after. | |
| 5 | 5 | //! | |
| 6 | 6 | //! The site reaches it over `POST /rpc/<method>` with the arguments below; | |
| 7 | 7 | //! the registries' own protocols are any other request. Mirrors | |
| 27 | 27 | Composer, | |
| 28 | 28 | Cargo, | |
| 29 | 29 | Go, | |
| 30 | + | Maven, | |
| 31 | + | Nuget, | |
| 32 | + | Rubygems, | |
| 30 | 33 | } | |
| 31 | 34 | ||
| 32 | 35 | impl Ecosystem { | |
| 33 | − | pub const ALL: [Ecosystem; 5] = [ | |
| 36 | + | pub const ALL: [Ecosystem; 8] = [ | |
| 34 | 37 | Ecosystem::Container, | |
| 35 | 38 | Ecosystem::Npm, | |
| 36 | 39 | Ecosystem::Composer, | |
| 37 | 40 | Ecosystem::Cargo, | |
| 38 | 41 | Ecosystem::Go, | |
| 42 | + | Ecosystem::Maven, | |
| 43 | + | Ecosystem::Nuget, | |
| 44 | + | Ecosystem::Rubygems, | |
| 39 | 45 | ]; | |
| 40 | 46 | ||
| 41 | 47 | pub fn as_str(self) -> &'static str { | |
| 45 | 51 | Ecosystem::Composer => "composer", | |
| 46 | 52 | Ecosystem::Cargo => "cargo", | |
| 47 | 53 | Ecosystem::Go => "go", | |
| 54 | + | Ecosystem::Maven => "maven", | |
| 55 | + | Ecosystem::Nuget => "nuget", | |
| 56 | + | Ecosystem::Rubygems => "rubygems", | |
| 48 | 57 | } | |
| 49 | 58 | } | |
| 50 | 59 |
| 21 | 21 | - **Content-addressed.** Every file is stored once by its SHA-256 (`blobs/sha256/<hex>`). A | |
| 22 | 22 | version is a list of files by digest. Pushing a layer two images share stores it once. | |
| 23 | 23 | - **The registry speaks each tool's own protocol**, unchanged. `docker`, `npm`, `composer`, | |
| 24 | − | `cargo` and `go` work with only a login and an address. | |
| 24 | + | `cargo`, `go`, `mvn` and Gradle, `dotnet` and `gem` and Bundler work with only a login and an | |
| 25 | + | address. | |
| 25 | 26 | - **Same access as the code.** A package can be linked to a repository and then has its | |
| 26 | 27 | visibility and roles. Unlinked, it is the workspace's: members by the base permission. | |
| 27 | 28 | - **Same front door.** Everything is on `g1t.sh`. `apps/web/workers/app.ts` hands registry paths | |
| 31 | 32 | ||
| 32 | 33 | | Table | What it holds | | |
| 33 | 34 | | --- | --- | | |
| 34 | − | | `packages` | `id`, `workspace`, `ecosystem` (`container`, `npm`, `composer`, `cargo`, `go`, ...), `name` (normalized per ecosystem), `repo_id` (linked repository, or null), `visibility` (`public`, `private`; linked packages follow their repository), `description`, `readme_digest`, `created_by`, `created_at`, `updated_at`, `downloads` | | |
| 35 | + | | `packages` | `id`, `workspace`, `ecosystem` (`container`, `npm`, `composer`, `cargo`, `go`, `maven`, `nuget`, `rubygems`), `name` (normalized per ecosystem), `repo_id` (linked repository, or null), `visibility` (`public`, `private`; linked packages follow their repository), `description`, `readme_digest`, `created_by`, `created_at`, `updated_at`, `downloads` | | |
| 35 | 36 | | `versions` | `id`, `package_id`, `version` (tag, semver or digest), `digest` (the manifest's or the archive's), `size` (sum of its files), `metadata` (JSON the ecosystem needs: npm's packument entry, a crate's index line, composer.json), `published_by`, `published_at`, `yanked`, `deprecated` | | |
| 36 | 37 | | `version_files` | `version_id`, `name`, `digest`, `size`, `media_type` | | |
| 37 | 38 | | `blobs` | `digest`, `size`, `media_type`, `created_at` | | |
| 38 | 39 | | `workspace_blobs` | `workspace`, `digest`, `public` (any public package uses it): what a workspace stores, counted once each, for billing | | |
| 39 | 40 | | `uploads` | `id`, `workspace`, `package`, `multipart_id`, `parts` (JSON), `offset`, `hash_state`, `expires_at`: uploads in progress | | |
| 40 | 41 | | `tags` | `package_id`, `tag`, `version_id`: container tags and npm dist-tags | | |
| 42 | + | | `checksums` | `digest`, `md5`, `sha1`, `sha512`: a file's other checksums, worked out on upload (Maven asks for them) | | |
| 41 | 43 | ||
| 42 | 44 | Deleting a version removes its rows; a daily sweep deletes blobs no version references, after a | |
| 43 | 45 | day's grace (a push in flight may reference a blob before its manifest lands). | |
| 119 | 121 | - A module proxy at `https://g1t.sh/-/go/` (`@v/list`, `.info`, `.mod`, `.zip`) built from tags, | |
| 120 | 122 | for faster and repeatable private installs (later phase). | |
| 121 | 123 | ||
| 124 | + | ### Maven | |
| 125 | + | ||
| 126 | + | - Per workspace: `https://g1t.sh/-/maven/<workspace>/`, the standard layout | |
| 127 | + | (`com/acme/web/1.0.0/web-1.0.0.jar`). A package is an artifact, named | |
| 128 | + | `groupId:artifactId`; a version holds every file uploaded into its | |
| 129 | + | directory, by file name (`version_files`), added one `PUT` at a time. | |
| 130 | + | - `mvn deploy` and Gradle's `publish`: Basic auth (any username, a token as | |
| 131 | + | the password) or `Bearer`. A release's files are written once (`409` for | |
| 132 | + | other content, the same content again is accepted); a SNAPSHOT's builds | |
| 133 | + | arrive as timestamped files beside each other. | |
| 134 | + | - Checksums: each file's MD5, SHA-1 and SHA-512 are worked out on upload | |
| 135 | + | and kept by digest (`checksums`); `.md5`, `.sha1`, `.sha256`, `.sha512` | |
| 136 | + | are answered from them, and uploaded ones are checked, not kept. | |
| 137 | + | - `maven-metadata.xml` is made on every read: per artifact (versions in | |
| 138 | + | Maven's order, `latest`, `release`) and per SNAPSHOT version (the newest | |
| 139 | + | build of each classifier and extension). Uploaded ones are accepted and | |
| 140 | + | let go, a plugin group's included. | |
| 141 | + | - The POM is the version's record: its coordinates must | |
| 142 | + | match its path, its description becomes the package's for the highest | |
| 143 | + | release, and on a new artifact its `<scm><url>` may link the repository. | |
| 144 | + | - The artifact's own `maven-metadata.xml`, uploaded last by Maven and Gradle, | |
| 145 | + | publishes (event, audit entry) each version or SNAPSHOT build whose POM | |
| 146 | + | arrived since, marked in its metadata so it is published once. | |
| 147 | + | ||
| 148 | + | ### NuGet | |
| 149 | + | ||
| 150 | + | - Per workspace: `https://g1t.sh/-/nuget/<workspace>/v3/index.json` naming | |
| 151 | + | `PackageBaseAddress/3.0.0` (flat container), `RegistrationsBaseUrl` | |
| 152 | + | (one inlined page), `SearchQueryService` and `PackagePublish/2.0.0`. | |
| 153 | + | - `dotnet nuget push`: a `PUT` of a multipart body with the `.nupkg`, | |
| 154 | + | `X-NuGet-ApiKey` a g1t token. The `.nuspec` (read from the zip) gives the | |
| 155 | + | id, version (normalized as NuGet does), description, dependency groups and | |
| 156 | + | README; the `.nupkg` and `.nuspec` are the version's files. | |
| 157 | + | - Ids are one whatever their case; a version is pushed once (`409`). | |
| 158 | + | - `DELETE api/v2/package/<id>/<version>` unlists (the `yanked` column), as | |
| 159 | + | nuget.org does; `POST` lists again. Unlisted versions stay in the flat | |
| 160 | + | container and registration (`listed: false`), not in search. | |
| 161 | + | - Restores use Basic auth from `nuget.config`, after a `401`. | |
| 162 | + | ||
| 163 | + | ### RubyGems | |
| 164 | + | ||
| 165 | + | - Per workspace: `https://g1t.sh/-/rubygems/<workspace>/`. `gem push` | |
| 166 | + | (`POST /api/v1/gems`, the token as the whole `Authorization` header), | |
| 167 | + | `gem yank` (`DELETE /api/v1/gems/yank`), downloads at | |
| 168 | + | `/gems/<name>-<version>[-<platform>].gem`. | |
| 169 | + | - The compact index Bundler reads: `versions`, `info/<gem>`, `names`, made | |
| 170 | + | on every read, each with the quoted MD5 of its body as its `ETag` (Bundler | |
| 171 | + | checks it, and `versions` names each info file's MD5). Yanked versions | |
| 172 | + | leave the index; their files stay. | |
| 173 | + | - The gem's `metadata.gz` (YAML, in the `.gem` tar) gives the name, | |
| 174 | + | version, platform and runtime dependencies. A version is keyed as the | |
| 175 | + | index writes it (`1.0.0`, `1.0.0-x86_64-linux`) and pushed once. | |
| 176 | + | - Bundler authenticates with Basic auth from `bundle config`. The full | |
| 177 | + | index (`specs.4.8.gz`, Marshal) is not served, so `gem install --source` | |
| 178 | + | is not supported. | |
| 179 | + | ||
| 122 | 180 | ### Later | |
| 123 | 181 | ||
| 124 | − | Maven (`/-/maven/<workspace>/`), NuGet (v3), RubyGems; PyPI. | |
| 182 | + | PyPI. | |
| 125 | 183 | ||
| 126 | 184 | ## Billing | |
| 127 | 185 | ||
| 162 | 220 | 2. **npm.** | |
| 163 | 221 | 3. **Composer and Go** (both from the repositories themselves). | |
| 164 | 222 | 4. **Cargo.** | |
| 165 | − | 5. **Mirrors** (Packagist, npm), **Maven, NuGet, RubyGems.** | |
| 223 | + | 5. **Maven, NuGet, RubyGems.** | |
| 224 | + | 6. **Mirrors** (Packagist, npm). |
| 13 | 13 | */ | |
| 14 | 14 | ||
| 15 | 15 | /** Which registry a package is in. */ | |
| 16 | − | export const ECOSYSTEMS = ["container", "npm", "composer", "cargo", "go"] as const; | |
| 16 | + | export const ECOSYSTEMS = ["container", "npm", "composer", "cargo", "go", "maven", "nuget", "rubygems"] as const; | |
| 17 | 17 | export type Ecosystem = (typeof ECOSYSTEMS)[number]; | |
| 18 | 18 | ||
| 19 | 19 | export type PackageVisibility = "public" | "private"; |
| 3 | 3 | version = "0.1.0" | |
| 4 | 4 | edition.workspace = true | |
| 5 | 5 | license.workspace = true | |
| 6 | − | description = "Package registries beside the code: container images over OCI Distribution, then npm, Composer, Cargo and Go." | |
| 6 | + | description = "Package registries beside the code: container images over OCI Distribution, npm, Composer, Cargo, Go, Maven, NuGet and RubyGems." | |
| 7 | 7 | ||
| 8 | 8 | [lib] | |
| 9 | 9 | crate-type = ["cdylib"] | |
| 20 | 20 | hmac = "0.12" | |
| 21 | 21 | sha2 = { version = "0.10", features = ["compress"] } | |
| 22 | 22 | sha1 = "0.10" | |
| 23 | + | md5 = "0.8" | |
| 23 | 24 | miniz_oxide = "0.9" | |
| 24 | 25 | ||
| 25 | 26 | # wasm-opt at -O1: about the same gzipped size as -O in a tenth of the |
| 1 | + | -- The checksums Maven asks for beside every file (`.md5`, `.sha1`, | |
| 2 | + | -- `.sha512`; `.sha256` is the blob's digest), worked out once when the | |
| 3 | + | -- file is uploaded so a download never reads a file twice. Kept by digest, | |
| 4 | + | -- as blobs are, and let go with them. | |
| 5 | + | CREATE TABLE checksums ( | |
| 6 | + | digest TEXT PRIMARY KEY, | |
| 7 | + | md5 TEXT NOT NULL, | |
| 8 | + | sha1 TEXT NOT NULL, | |
| 9 | + | sha512 TEXT NOT NULL | |
| 10 | + | ); |
| 1 | + | //! Reading the archives packages arrive as: a `.nupkg` is a zip (read with | |
| 2 | + | //! the CRC-32 the Composer zips are written with), a `.gem` is a tar | |
| 3 | + | //! holding gzipped files. Only what a registry needs is read: the entries' | |
| 4 | + | //! names, and the bytes of the few it asks for, each up to a limit. | |
| 5 | + | ||
| 6 | + | use crate::composer::crc32; | |
| 7 | + | ||
| 8 | + | /// One file in a zip, as its central directory lists it. | |
| 9 | + | #[derive(Clone, Debug, PartialEq, Eq)] | |
| 10 | + | pub struct ZipEntry { | |
| 11 | + | pub name: String, | |
| 12 | + | method: u16, | |
| 13 | + | crc: u32, | |
| 14 | + | compressed: u64, | |
| 15 | + | pub size: u64, | |
| 16 | + | offset: u64, | |
| 17 | + | } | |
| 18 | + | ||
| 19 | + | fn u16_at(bytes: &[u8], at: usize) -> Option<u16> { | |
| 20 | + | Some(u16::from_le_bytes(bytes.get(at..at + 2)?.try_into().ok()?)) | |
| 21 | + | } | |
| 22 | + | ||
| 23 | + | fn u32_at(bytes: &[u8], at: usize) -> Option<u32> { | |
| 24 | + | Some(u32::from_le_bytes(bytes.get(at..at + 4)?.try_into().ok()?)) | |
| 25 | + | } | |
| 26 | + | ||
| 27 | + | /// The files a zip holds, from its central directory. | |
| 28 | + | pub fn zip_entries(bytes: &[u8]) -> Result<Vec<ZipEntry>, String> { | |
| 29 | + | const END: u32 = 0x0605_4b50; | |
| 30 | + | const CENTRAL: u32 = 0x0201_4b50; | |
| 31 | + | let not_zip = || "The file is not a zip archive.".to_owned(); | |
| 32 | + | if bytes.len() < 22 { | |
| 33 | + | return Err(not_zip()); | |
| 34 | + | } | |
| 35 | + | // The end record is the last 22 bytes, before a comment of up to 64 KB. | |
| 36 | + | let earliest = bytes.len().saturating_sub(22 + 0xFFFF); | |
| 37 | + | let end = (earliest..=bytes.len() - 22).rev().find(|&at| u32_at(bytes, at) == Some(END)).ok_or_else(not_zip)?; | |
| 38 | + | let count = u16_at(bytes, end + 10).ok_or_else(not_zip)?; | |
| 39 | + | let mut at = u32_at(bytes, end + 16).ok_or_else(not_zip)? as usize; | |
| 40 | + | if count == 0xFFFF || at == 0xFFFF_FFFF_usize { | |
| 41 | + | return Err("The archive is a zip64 archive, which is not read here.".to_owned()); | |
| 42 | + | } | |
| 43 | + | let mut entries = Vec::with_capacity(count as usize); | |
| 44 | + | for _ in 0..count { | |
| 45 | + | if u32_at(bytes, at) != Some(CENTRAL) { | |
| 46 | + | return Err("The zip's directory is damaged.".to_owned()); | |
| 47 | + | } | |
| 48 | + | let field16 = |offset: usize| u16_at(bytes, at + offset).ok_or_else(not_zip); | |
| 49 | + | let field32 = |offset: usize| u32_at(bytes, at + offset).ok_or_else(not_zip); | |
| 50 | + | let (name_len, extra_len, comment_len) = (field16(28)? as usize, field16(30)? as usize, field16(32)? as usize); | |
| 51 | + | let name = bytes.get(at + 46..at + 46 + name_len).ok_or_else(not_zip)?; | |
| 52 | + | entries.push(ZipEntry { | |
| 53 | + | name: String::from_utf8_lossy(name).into_owned(), | |
| 54 | + | method: field16(10)?, | |
| 55 | + | crc: field32(16)?, | |
| 56 | + | compressed: u64::from(field32(20)?), | |
| 57 | + | size: u64::from(field32(24)?), | |
| 58 | + | offset: u64::from(field32(42)?), | |
| 59 | + | }); | |
| 60 | + | at += 46 + name_len + extra_len + comment_len; | |
| 61 | + | } | |
| 62 | + | Ok(entries) | |
| 63 | + | } | |
| 64 | + | ||
| 65 | + | /// The bytes of one entry, inflated and checked against its CRC-32. An | |
| 66 | + | /// entry larger than `limit` is refused. | |
| 67 | + | pub fn zip_read(bytes: &[u8], entry: &ZipEntry, limit: usize) -> Result<Vec<u8>, String> { | |
| 68 | + | const LOCAL: u32 = 0x0403_4b50; | |
| 69 | + | let damaged = || format!("{} is damaged in the archive.", entry.name); | |
| 70 | + | if entry.size > limit as u64 { | |
| 71 | + | return Err(format!("{} is larger than {} KB.", entry.name, limit / 1024)); | |
| 72 | + | } | |
| 73 | + | let at = entry.offset as usize; | |
| 74 | + | if u32_at(bytes, at) != Some(LOCAL) { | |
| 75 | + | return Err(damaged()); | |
| 76 | + | } | |
| 77 | + | let start = at + 30 + u16_at(bytes, at + 26).ok_or_else(damaged)? as usize + u16_at(bytes, at + 28).ok_or_else(damaged)? as usize; | |
| 78 | + | let body = bytes.get(start..start + entry.compressed as usize).ok_or_else(damaged)?; | |
| 79 | + | let data = match entry.method { | |
| 80 | + | 0 => body.to_vec(), | |
| 81 | + | 8 => miniz_oxide::inflate::decompress_to_vec_with_limit(body, limit).map_err(|_| damaged())?, | |
| 82 | + | other => return Err(format!("{} is compressed with method {other}, which is not read here.", entry.name)), | |
| 83 | + | }; | |
| 84 | + | if data.len() as u64 != entry.size || crc32(&data) != entry.crc { | |
| 85 | + | return Err(damaged()); | |
| 86 | + | } | |
| 87 | + | Ok(data) | |
| 88 | + | } | |
| 89 | + | ||
| 90 | + | /// The bytes of a gzip file, inflated, up to `limit`. | |
| 91 | + | pub fn gunzip(bytes: &[u8], limit: usize) -> Result<Vec<u8>, String> { | |
| 92 | + | let bad = || "The file is not gzipped.".to_owned(); | |
| 93 | + | if bytes.len() < 18 || bytes[0] != 0x1f || bytes[1] != 0x8b || bytes[2] != 8 { | |
| 94 | + | return Err(bad()); | |
| 95 | + | } | |
| 96 | + | let flags = bytes[3]; | |
| 97 | + | let mut at = 10; | |
| 98 | + | if flags & 4 != 0 { | |
| 99 | + | at += 2 + u16_at(bytes, at).ok_or_else(bad)? as usize; | |
| 100 | + | } | |
| 101 | + | for flag in [8u8, 16] { | |
| 102 | + | if flags & flag != 0 { | |
| 103 | + | at += bytes.get(at..).ok_or_else(bad)?.iter().position(|b| *b == 0).ok_or_else(bad)? + 1; | |
| 104 | + | } | |
| 105 | + | } | |
| 106 | + | if flags & 2 != 0 { | |
| 107 | + | at += 2; | |
| 108 | + | } | |
| 109 | + | let body = bytes.get(at..bytes.len() - 8).ok_or_else(bad)?; | |
| 110 | + | miniz_oxide::inflate::decompress_to_vec_with_limit(body, limit).map_err(|_| "The gzipped file is damaged or too large.".to_owned()) | |
| 111 | + | } | |
| 112 | + | ||
| 113 | + | /// The regular files of a tar, as name and bytes. | |
| 114 | + | pub fn tar_files(bytes: &[u8]) -> Result<Vec<(String, &[u8])>, String> { | |
| 115 | + | let mut files = Vec::new(); | |
| 116 | + | let mut at = 0; | |
| 117 | + | while at + 512 <= bytes.len() { | |
| 118 | + | let header = &bytes[at..at + 512]; | |
| 119 | + | if header.iter().all(|b| *b == 0) { | |
| 120 | + | break; | |
| 121 | + | } | |
| 122 | + | let text = |range: std::ops::Range<usize>| { | |
| 123 | + | let field = &header[range]; | |
| 124 | + | let end = field.iter().position(|b| *b == 0).unwrap_or(field.len()); | |
| 125 | + | String::from_utf8_lossy(&field[..end]).into_owned() | |
| 126 | + | }; | |
| 127 | + | let size = u64::from_str_radix(text(124..136).trim(), 8).map_err(|_| "The tar's header is damaged.".to_owned())? as usize; | |
| 128 | + | let mut name = text(0..100); | |
| 129 | + | if &header[257..262] == b"ustar" { | |
| 130 | + | let prefix = text(345..500); | |
| 131 | + | if !prefix.is_empty() { | |
| 132 | + | name = format!("{prefix}/{name}"); | |
| 133 | + | } | |
| 134 | + | } | |
| 135 | + | let start = at + 512; | |
| 136 | + | let data = bytes.get(start..start + size).ok_or("The tar ends early.")?; | |
| 137 | + | if matches!(header[156], 0 | b'0') { | |
| 138 | + | files.push((name, data)); | |
| 139 | + | } | |
| 140 | + | at = start + size.div_ceil(512) * 512; | |
| 141 | + | } | |
| 142 | + | Ok(files) | |
| 143 | + | } | |
| 144 | + | ||
| 145 | + | #[cfg(test)] | |
| 146 | + | mod tests { | |
| 147 | + | use super::*; | |
| 148 | + | ||
| 149 | + | #[test] | |
| 150 | + | fn a_zip_written_here_reads_back() { | |
| 151 | + | let files = vec![ | |
| 152 | + | ("Acme.Web.nuspec".to_owned(), b"<package/>".to_vec()), | |
| 153 | + | ("lib/net8.0/Acme.Web.dll".to_owned(), "MZ".repeat(500).into_bytes()), | |
| 154 | + | ]; | |
| 155 | + | let zip = crate::composer::zip(&files); | |
| 156 | + | let entries = zip_entries(&zip).unwrap(); | |
| 157 | + | assert_eq!(entries.iter().map(|e| e.name.as_str()).collect::<Vec<_>>(), ["Acme.Web.nuspec", "lib/net8.0/Acme.Web.dll"]); | |
| 158 | + | assert_eq!(zip_read(&zip, &entries[0], 1024).unwrap(), b"<package/>"); | |
| 159 | + | assert_eq!(zip_read(&zip, &entries[1], 4096).unwrap(), "MZ".repeat(500).into_bytes(), "deflated"); | |
| 160 | + | assert!(zip_read(&zip, &entries[1], 100).is_err(), "over the limit"); | |
| 161 | + | let mut damaged = zip.clone(); | |
| 162 | + | damaged[30 + "Acme.Web.nuspec".len() + 2] ^= 0xFF; | |
| 163 | + | assert!(zip_read(&damaged, &entries[0], 1024).is_err(), "the CRC catches it"); | |
| 164 | + | assert!(zip_entries(b"not a zip at all, but long enough").is_err()); | |
| 165 | + | } | |
| 166 | + | ||
| 167 | + | fn gzip(data: &[u8]) -> Vec<u8> { | |
| 168 | + | let mut out = vec![0x1f, 0x8b, 8, 8, 0, 0, 0, 0, 0, 3]; | |
| 169 | + | out.extend_from_slice(b"metadata\0"); | |
| 170 | + | out.extend_from_slice(&miniz_oxide::deflate::compress_to_vec(data, 6)); | |
| 171 | + | out.extend_from_slice(&crc32(data).to_le_bytes()); | |
| 172 | + | out.extend_from_slice(&(data.len() as u32).to_le_bytes()); | |
| 173 | + | out | |
| 174 | + | } | |
| 175 | + | ||
| 176 | + | pub fn tar(files: &[(&str, &[u8])]) -> Vec<u8> { | |
| 177 | + | let mut out = Vec::new(); | |
| 178 | + | for (name, data) in files { | |
| 179 | + | let mut header = [0u8; 512]; | |
| 180 | + | header[..name.len()].copy_from_slice(name.as_bytes()); | |
| 181 | + | let size = format!("{:011o}\0", data.len()); | |
| 182 | + | header[124..136].copy_from_slice(size.as_bytes()); | |
| 183 | + | header[156] = b'0'; | |
| 184 | + | header[257..262].copy_from_slice(b"ustar"); | |
| 185 | + | out.extend_from_slice(&header); | |
| 186 | + | out.extend_from_slice(data); | |
| 187 | + | out.resize(out.len().div_ceil(512) * 512, 0); | |
| 188 | + | } | |
| 189 | + | out.extend_from_slice(&[0; 1024]); | |
| 190 | + | out | |
| 191 | + | } | |
| 192 | + | ||
| 193 | + | #[test] | |
| 194 | + | fn a_gem_is_a_tar_of_gzipped_files() { | |
| 195 | + | let metadata = gzip(b"--- !ruby/object:Gem::Specification\nname: hello\n"); | |
| 196 | + | let gem = tar(&[("metadata.gz", &metadata), ("data.tar.gz", b"data")]); | |
| 197 | + | let files = tar_files(&gem).unwrap(); | |
| 198 | + | assert_eq!(files.len(), 2); | |
| 199 | + | assert_eq!(files[0].0, "metadata.gz"); | |
| 200 | + | assert_eq!(files[1].1, b"data"); | |
| 201 | + | assert_eq!(gunzip(files[0].1, 1024).unwrap(), b"--- !ruby/object:Gem::Specification\nname: hello\n"); | |
| 202 | + | assert!(gunzip(b"plain text, not gzip at all", 1024).is_err()); | |
| 203 | + | assert!(tar_files(&gem[..1538]).is_err(), "cut short"); | |
| 204 | + | } | |
| 205 | + | } |
| 283 | 283 | } | |
| 284 | 284 | ||
| 285 | 285 | /// CRC-32, as zip records each file's. | |
| 286 | − | fn crc32(bytes: &[u8]) -> u32 { | |
| 286 | + | pub(crate) fn crc32(bytes: &[u8]) -> u32 { | |
| 287 | 287 | let mut table = [0u32; 256]; | |
| 288 | 288 | for (i, entry) in table.iter_mut().enumerate() { | |
| 289 | 289 | let mut c = i as u32; |
| 240 | 240 | } | |
| 241 | 241 | } | |
| 242 | 242 | ||
| 243 | + | /// One file of a version, as it is kept. | |
| 244 | + | #[derive(Clone, Debug, Deserialize)] | |
| 245 | + | pub struct FileRow { | |
| 246 | + | pub name: String, | |
| 247 | + | pub digest: String, | |
| 248 | + | pub size: u64, | |
| 249 | + | pub media_type: Option<String>, | |
| 250 | + | } | |
| 251 | + | ||
| 252 | + | /// A file's other checksums, in hex, beside its SHA-256 digest. | |
| 253 | + | #[derive(Clone, Debug, PartialEq, Eq, Deserialize)] | |
| 254 | + | pub struct Checksums { | |
| 255 | + | pub md5: String, | |
| 256 | + | pub sha1: String, | |
| 257 | + | pub sha512: String, | |
| 258 | + | } | |
| 259 | + | ||
| 260 | + | impl Checksums { | |
| 261 | + | pub fn of(bytes: &[u8]) -> Checksums { | |
| 262 | + | use sha1::Digest as _; | |
| 263 | + | Checksums { | |
| 264 | + | md5: format!("{:x}", md5::compute(bytes)), | |
| 265 | + | sha1: hex::encode(sha1::Sha1::digest(bytes)), | |
| 266 | + | sha512: hex::encode(sha2::Sha512::digest(bytes)), | |
| 267 | + | } | |
| 268 | + | } | |
| 269 | + | } | |
| 270 | + | ||
| 243 | 271 | /// One file of a version, as it is recorded. | |
| 244 | 272 | pub struct NewFile { | |
| 245 | 273 | pub name: String, | |
| 1045 | 1073 | .results() | |
| 1046 | 1074 | } | |
| 1047 | 1075 | ||
| 1076 | + | /// A version by its version string, made from `version` when there is | |
| 1077 | + | /// none, and whether it was made now. Files are added one at a time | |
| 1078 | + | /// with `put_file` (Maven uploads each in a request of its own). | |
| 1079 | + | pub async fn version_or_new(&self, version: &NewVersion, now_ms: u64) -> Result<(VersionRow, bool)> { | |
| 1080 | + | self.prepare( | |
| 1081 | + | "INSERT OR IGNORE INTO versions (id, package_id, version, digest, size, metadata, subject, published_by, published_at) | |
| 1082 | + | VALUES (?, ?, ?, ?, 0, ?, NULL, ?, ?)", | |
| 1083 | + | &[ | |
| 1084 | + | text(&version.id), | |
| 1085 | + | text(&version.package_id), | |
| 1086 | + | text(&version.version), | |
| 1087 | + | text(&version.digest), | |
| 1088 | + | text(&version.metadata), | |
| 1089 | + | opt(version.published_by.as_deref()), | |
| 1090 | + | text(&rfc3339(now_ms)), | |
| 1091 | + | ], | |
| 1092 | + | )? | |
| 1093 | + | .run() | |
| 1094 | + | .await?; | |
| 1095 | + | let row = self | |
| 1096 | + | .version_named(&version.package_id, &version.version) | |
| 1097 | + | .await? | |
| 1098 | + | .ok_or_else(|| worker::Error::RustError("the version was not recorded".into()))?; | |
| 1099 | + | let made = row.id == version.id; | |
| 1100 | + | Ok((row, made)) | |
| 1101 | + | } | |
| 1102 | + | ||
| 1103 | + | /// Adds a file to a version, or replaces the one of its name, and | |
| 1104 | + | /// works out the version's size again. | |
| 1105 | + | pub async fn put_file(&self, package_id: &str, version_id: &str, file: &NewFile, now_ms: u64) -> Result<()> { | |
| 1106 | + | let now = rfc3339(now_ms); | |
| 1107 | + | self.db | |
| 1108 | + | .batch(vec![ | |
| 1109 | + | self.prepare( | |
| 1110 | + | "INSERT INTO version_files (version_id, name, digest, size, media_type) VALUES (?, ?, ?, ?, ?) | |
| 1111 | + | ON CONFLICT (version_id, name) DO UPDATE SET digest = excluded.digest, size = excluded.size, media_type = excluded.media_type", | |
| 1112 | + | &[text(version_id), text(&file.name), text(&file.digest), num(file.size), opt(file.media_type.as_deref())], | |
| 1113 | + | )?, | |
| 1114 | + | self.prepare( | |
| 1115 | + | "UPDATE versions SET size = (SELECT COALESCE(SUM(size), 0) FROM version_files WHERE version_id = ?) WHERE id = ?", | |
| 1116 | + | &[text(version_id), text(version_id)], | |
| 1117 | + | )?, | |
| 1118 | + | self.prepare("UPDATE packages SET updated_at = ? WHERE id = ?", &[text(&now), text(package_id)])?, | |
| 1119 | + | ]) | |
| 1120 | + | .await?; | |
| 1121 | + | Ok(()) | |
| 1122 | + | } | |
| 1123 | + | ||
| 1124 | + | /// A version's files, by name. | |
| 1125 | + | pub async fn files(&self, version_id: &str) -> Result<Vec<FileRow>> { | |
| 1126 | + | self.prepare("SELECT name, digest, size, media_type FROM version_files WHERE version_id = ? ORDER BY name", &[text(version_id)])? | |
| 1127 | + | .all() | |
| 1128 | + | .await? | |
| 1129 | + | .results() | |
| 1130 | + | } | |
| 1131 | + | ||
| 1132 | + | pub async fn file(&self, version_id: &str, name: &str) -> Result<Option<FileRow>> { | |
| 1133 | + | self.prepare( | |
| 1134 | + | "SELECT name, digest, size, media_type FROM version_files WHERE version_id = ? AND name = ?", | |
| 1135 | + | &[text(version_id), text(name)], | |
| 1136 | + | )? | |
| 1137 | + | .first(None) | |
| 1138 | + | .await | |
| 1139 | + | } | |
| 1140 | + | ||
| 1141 | + | /// Changes what a version is named by and keeps about itself. | |
| 1142 | + | pub async fn set_version(&self, version_id: &str, digest: &str, metadata: &str) -> Result<()> { | |
| 1143 | + | self.prepare("UPDATE versions SET digest = ?, metadata = ? WHERE id = ?", &[text(digest), text(metadata), text(version_id)])? | |
| 1144 | + | .run() | |
| 1145 | + | .await?; | |
| 1146 | + | Ok(()) | |
| 1147 | + | } | |
| 1148 | + | ||
| 1149 | + | pub async fn set_checksums(&self, digest: &Digest, sums: &Checksums) -> Result<()> { | |
| 1150 | + | self.prepare( | |
| 1151 | + | "INSERT OR IGNORE INTO checksums (digest, md5, sha1, sha512) VALUES (?, ?, ?, ?)", | |
| 1152 | + | &[text(digest.as_str()), text(&sums.md5), text(&sums.sha1), text(&sums.sha512)], | |
| 1153 | + | )? | |
| 1154 | + | .run() | |
| 1155 | + | .await?; | |
| 1156 | + | Ok(()) | |
| 1157 | + | } | |
| 1158 | + | ||
| 1159 | + | pub async fn checksums(&self, digest: &Digest) -> Result<Option<Checksums>> { | |
| 1160 | + | self.prepare("SELECT md5, sha1, sha512 FROM checksums WHERE digest = ?", &[text(digest.as_str())])? | |
| 1161 | + | .first(None) | |
| 1162 | + | .await | |
| 1163 | + | } | |
| 1164 | + | ||
| 1165 | + | /// Every version of a workspace's packages of an ecosystem, oldest | |
| 1166 | + | /// first: what an index of the whole registry is made from. | |
| 1167 | + | pub async fn ecosystem_versions(&self, workspace: &str, ecosystem: &str, limit: u32) -> Result<Vec<VersionRow>> { | |
| 1168 | + | self.prepare( | |
| 1169 | + | &format!( | |
| 1170 | + | "SELECT {} FROM versions v JOIN packages p ON p.id = v.package_id | |
| 1171 | + | WHERE p.workspace = ? AND p.ecosystem = ? AND p.workspace_deleted_at IS NULL | |
| 1172 | + | ORDER BY v.published_at, v.id LIMIT {limit}", | |
| 1173 | + | VERSION_COLUMNS.split(", ").map(|c| format!("v.{c}")).collect::<Vec<_>>().join(", ") | |
| 1174 | + | ), | |
| 1175 | + | &[text(workspace), text(ecosystem)], | |
| 1176 | + | )? | |
| 1177 | + | .all() | |
| 1178 | + | .await? | |
| 1179 | + | .results() | |
| 1180 | + | } | |
| 1181 | + | ||
| 1182 | + | /// A workspace's packages of an ecosystem, by name. | |
| 1183 | + | pub async fn packages_of(&self, workspace: &str, ecosystem: &str, limit: u32) -> Result<Vec<PackageRow>> { | |
| 1184 | + | self.prepare( | |
| 1185 | + | &format!( | |
| 1186 | + | "SELECT {PACKAGE_COLUMNS} FROM packages WHERE workspace = ? AND ecosystem = ? AND workspace_deleted_at IS NULL ORDER BY name LIMIT {limit}" | |
| 1187 | + | ), | |
| 1188 | + | &[text(workspace), text(ecosystem)], | |
| 1189 | + | )? | |
| 1190 | + | .all() | |
| 1191 | + | .await? | |
| 1192 | + | .results() | |
| 1193 | + | } | |
| 1194 | + | ||
| 1048 | 1195 | pub async fn forget_blob(&self, digest: &str) -> Result<()> { | |
| 1049 | 1196 | let d = [text(digest)]; | |
| 1050 | 1197 | self.db | |
| 1051 | 1198 | .batch(vec![ | |
| 1199 | + | self.prepare("DELETE FROM checksums WHERE digest = ?", &d)?, | |
| 1052 | 1200 | self.prepare("DELETE FROM package_blobs WHERE digest = ?", &d)?, | |
| 1053 | 1201 | self.prepare("DELETE FROM workspace_blobs WHERE digest = ?", &d)?, | |
| 1054 | 1202 | self.prepare("DELETE FROM blobs WHERE digest = ?", &d)?, |
| 8 | 8 | //! `BlobStore` port (store/), metadata in D1 (db.rs). | |
| 9 | 9 | ||
| 10 | 10 | mod access; | |
| 11 | + | mod archive; | |
| 11 | 12 | mod cargo; | |
| 12 | 13 | mod cargo_http; | |
| 13 | 14 | mod composer; | |
| 16 | 17 | mod digest; | |
| 17 | 18 | mod limits; | |
| 18 | 19 | mod manifest; | |
| 20 | + | mod maven; | |
| 21 | + | mod maven_http; | |
| 19 | 22 | mod names; | |
| 20 | 23 | mod npm; | |
| 21 | 24 | mod npm_http; | |
| 25 | + | mod nuget; | |
| 26 | + | mod nuget_http; | |
| 22 | 27 | mod oci; | |
| 23 | 28 | mod quota; | |
| 24 | 29 | mod range; | |
| 30 | + | mod rubygems; | |
| 31 | + | mod rubygems_http; | |
| 25 | 32 | mod sigv4; | |
| 26 | 33 | mod store; | |
| 27 | 34 | mod token; | |
| 28 | 35 | mod upload; | |
| 36 | + | mod xml; | |
| 37 | + | mod yaml; | |
| 29 | 38 | ||
| 30 | 39 | use std::cell::RefCell; | |
| 31 | 40 | use std::collections::HashMap; | |
| 351 | 360 | "npm" => format!("{}/-/npm/@{}/{}", self.host, p.workspace, p.name), | |
| 352 | 361 | "composer" => format!("{}/-/composer/{}/{}", self.host, p.workspace, p.name), | |
| 353 | 362 | "cargo" => format!("{}/-/cargo/{}/{}", self.host, p.workspace, p.name), | |
| 363 | + | "maven" => format!("{}/-/maven/{}/{}", self.host, p.workspace, p.name), | |
| 364 | + | "nuget" => format!("{}/-/nuget/{}/{}", self.host, p.workspace, p.name), | |
| 365 | + | "rubygems" => format!("{}/-/rubygems/{}/{}", self.host, p.workspace, p.name), | |
| 354 | 366 | _ => format!("{}/{}/{}", self.host, p.workspace, p.name), | |
| 355 | 367 | }, | |
| 356 | 368 | visibility: Visibility::parse(&p.visibility), | |
| 410 | 422 | return Ok(not_found()); | |
| 411 | 423 | } | |
| 412 | 424 | let tags = self.db.tags(&row.package.id).await?; | |
| 425 | + | // A yanked or unlisted version reads as deprecated: still there | |
| 426 | + | // for lockfiles that name it, no longer picked for new ones. | |
| 427 | + | let withdrawn = match row.package.ecosystem.as_str() { | |
| 428 | + | "nuget" => "Unlisted: still restored by projects that name it, no longer shown in search.", | |
| 429 | + | "rubygems" => "Yanked: Bundler no longer picks this version for new lockfiles.", | |
| 430 | + | _ => "Yanked: Cargo no longer picks this version for new lockfiles.", | |
| 431 | + | }; | |
| 413 | 432 | let versions = self | |
| 414 | 433 | .db | |
| 415 | 434 | .versions(&row.package.id, VERSIONS_SHOWN) | |
| 433 | 452 | subject: version.subject, | |
| 434 | 453 | published_by: version.published_by, | |
| 435 | 454 | published_at: version.published_at, | |
| 436 | − | // A yanked crate version reads as deprecated: still | |
| 437 | − | // there for lockfiles, no longer picked for new ones. | |
| 438 | 455 | deprecated: if version.yanked != 0 { | |
| 439 | − | Some("Yanked: Cargo no longer picks this version for new lockfiles.".to_owned()) | |
| 456 | + | Some(withdrawn.to_owned()) | |
| 440 | 457 | } else { | |
| 441 | 458 | version.deprecated | |
| 442 | 459 | }, | |
| 682 | 699 | if request.path().starts_with("/-/cargo/") { | |
| 683 | 700 | return packages.cargo(request, &ctx).await; | |
| 684 | 701 | } | |
| 702 | + | if request.path().starts_with("/-/maven/") { | |
| 703 | + | return packages.maven(request, &ctx).await; | |
| 704 | + | } | |
| 705 | + | if request.path().starts_with("/-/nuget/") { | |
| 706 | + | return packages.nuget(request, &ctx).await; | |
| 707 | + | } | |
| 708 | + | if request.path().starts_with("/-/rubygems/") { | |
| 709 | + | return packages.rubygems(request, &ctx).await; | |
| 710 | + | } | |
| 685 | 711 | return packages.registry(request, &ctx).await; | |
| 686 | 712 | }; | |
| 687 | 713 | let body: serde_json::Value = request.json().await?; |
| 1 | + | //! What the Maven repository needs that does not touch the network: the | |
| 2 | + | //! standard layout's paths (`com/acme/web/1.0.0/web-1.0.0.jar`), the | |
| 3 | + | //! checksum files beside each file, the `maven-metadata.xml` made for each | |
| 4 | + | //! artifact and SNAPSHOT version, Maven's version order, and the few | |
| 5 | + | //! things read from a POM. | |
| 6 | + | //! | |
| 7 | + | //! A package is an artifact, named `groupId:artifactId`. A version holds | |
| 8 | + | //! every file uploaded into its directory, by file name; a SNAPSHOT | |
| 9 | + | //! version holds each build's timestamped files, and its metadata names | |
| 10 | + | //! the newest of each. | |
| 11 | + | ||
| 12 | + | use std::cmp::Ordering; | |
| 13 | + | ||
| 14 | + | use crate::xml; | |
| 15 | + | ||
| 16 | + | /// The checksum files Maven and Gradle upload and ask for beside a file. | |
| 17 | + | #[derive(Clone, Copy, Debug, PartialEq, Eq)] | |
| 18 | + | pub enum Checksum { | |
| 19 | + | Md5, | |
| 20 | + | Sha1, | |
| 21 | + | Sha256, | |
| 22 | + | Sha512, | |
| 23 | + | } | |
| 24 | + | ||
| 25 | + | impl Checksum { | |
| 26 | + | const ALL: [Checksum; 4] = [Checksum::Md5, Checksum::Sha1, Checksum::Sha256, Checksum::Sha512]; | |
| 27 | + | ||
| 28 | + | pub fn suffix(self) -> &'static str { | |
| 29 | + | match self { | |
| 30 | + | Checksum::Md5 => ".md5", | |
| 31 | + | Checksum::Sha1 => ".sha1", | |
| 32 | + | Checksum::Sha256 => ".sha256", | |
| 33 | + | Checksum::Sha512 => ".sha512", | |
| 34 | + | } | |
| 35 | + | } | |
| 36 | + | ||
| 37 | + | /// This checksum of `bytes`, in hex. | |
| 38 | + | pub fn of(self, bytes: &[u8]) -> String { | |
| 39 | + | use sha2::Digest as _; | |
| 40 | + | match self { | |
| 41 | + | Checksum::Md5 => format!("{:x}", md5::compute(bytes)), | |
| 42 | + | Checksum::Sha1 => hex::encode(sha1::Sha1::digest(bytes)), | |
| 43 | + | Checksum::Sha256 => hex::encode(sha2::Sha256::digest(bytes)), | |
| 44 | + | Checksum::Sha512 => hex::encode(sha2::Sha512::digest(bytes)), | |
| 45 | + | } | |
| 46 | + | } | |
| 47 | + | ||
| 48 | + | /// This checksum from the ones kept for a file, and its SHA-256 digest. | |
| 49 | + | pub fn pick(self, sums: &crate::db::Checksums, sha256: &str) -> String { | |
| 50 | + | match self { | |
| 51 | + | Checksum::Md5 => sums.md5.clone(), | |
| 52 | + | Checksum::Sha1 => sums.sha1.clone(), | |
| 53 | + | Checksum::Sha256 => sha256.to_owned(), | |
| 54 | + | Checksum::Sha512 => sums.sha512.clone(), | |
| 55 | + | } | |
| 56 | + | } | |
| 57 | + | } | |
| 58 | + | ||
| 59 | + | /// A file name without its checksum suffix, and which checksum it is. | |
| 60 | + | pub fn split_checksum(file: &str) -> (&str, Option<Checksum>) { | |
| 61 | + | for checksum in Checksum::ALL { | |
| 62 | + | if let Some(base) = file.strip_suffix(checksum.suffix()) { | |
| 63 | + | return (base, Some(checksum)); | |
| 64 | + | } | |
| 65 | + | } | |
| 66 | + | (file, None) | |
| 67 | + | } | |
| 68 | + | ||
| 69 | + | /// The hex a checksum file holds: its first word (some tools add the file | |
| 70 | + | /// name after it), in lowercase. | |
| 71 | + | pub fn sent_checksum(body: &[u8]) -> String { | |
| 72 | + | String::from_utf8_lossy(body).split_whitespace().next().unwrap_or("").to_ascii_lowercase() | |
| 73 | + | } | |
| 74 | + | ||
| 75 | + | const METADATA: &str = "maven-metadata.xml"; | |
| 76 | + | pub const MAX_PART: usize = 128; | |
| 77 | + | ||
| 78 | + | /// One directory of a groupId: `com` and `acme` of `com.acme`. | |
| 79 | + | fn valid_group_part(part: &str) -> bool { | |
| 80 | + | !part.is_empty() && part.len() <= MAX_PART && part.bytes().all(|b| b.is_ascii_alphanumeric() || matches!(b, b'-' | b'_')) | |
| 81 | + | } | |
| 82 | + | ||
| 83 | + | /// An artifactId: letters, digits, `-`, `_` and `.`, not starting with `.`. | |
| 84 | + | pub fn valid_artifact(artifact: &str) -> bool { | |
| 85 | + | !artifact.is_empty() | |
| 86 | + | && artifact.len() <= MAX_PART | |
| 87 | + | && !artifact.starts_with('.') | |
| 88 | + | && artifact.bytes().all(|b| b.is_ascii_alphanumeric() || matches!(b, b'-' | b'_' | b'.')) | |
| 89 | + | } | |
| 90 | + | ||
| 91 | + | /// A version: letters, digits, `.`, `-`, `_` and `+`, starting with a | |
| 92 | + | /// letter or digit. | |
| 93 | + | pub fn valid_version(version: &str) -> bool { | |
| 94 | + | !version.is_empty() | |
| 95 | + | && version.len() <= MAX_PART | |
| 96 | + | && version.as_bytes()[0].is_ascii_alphanumeric() | |
| 97 | + | && version.bytes().all(|b| b.is_ascii_alphanumeric() || matches!(b, b'-' | b'_' | b'.' | b'+')) | |
| 98 | + | && version != METADATA | |
| 99 | + | } | |
| 100 | + | ||
| 101 | + | pub fn is_snapshot(version: &str) -> bool { | |
| 102 | + | version.ends_with("-SNAPSHOT") | |
| 103 | + | } | |
| 104 | + | ||
| 105 | + | /// The package's name: `com.acme:web`. | |
| 106 | + | pub fn package_name(group: &str, artifact: &str) -> String { | |
| 107 | + | format!("{group}:{artifact}") | |
| 108 | + | } | |
| 109 | + | ||
| 110 | + | /// What a path under `/-/maven/<workspace>/` is. | |
| 111 | + | #[derive(Clone, Debug, PartialEq, Eq)] | |
| 112 | + | pub enum MavenPath { | |
| 113 | + | /// `com/acme/web/maven-metadata.xml`: an artifact's versions. | |
| 114 | + | ArtifactMetadata { group: String, artifact: String, checksum: Option<Checksum> }, | |
| 115 | + | /// `com/acme/web/1.0-SNAPSHOT/maven-metadata.xml`: a SNAPSHOT's builds. | |
| 116 | + | VersionMetadata { group: String, artifact: String, version: String, checksum: Option<Checksum> }, | |
| 117 | + | /// `com/acme/web/1.0.0/web-1.0.0.jar`: one of a version's files. | |
| 118 | + | File { group: String, artifact: String, version: String, file: String, checksum: Option<Checksum> }, | |
| 119 | + | } | |
| 120 | + | ||
| 121 | + | /// The workspace and what the path is. Names and versions are checked, | |
| 122 | + | /// and a file's name has to start with its artifactId. | |
| 123 | + | pub fn route(path: &str) -> Option<(String, MavenPath)> { | |
| 124 | + | let rest = path.strip_prefix("/-/maven/")?; | |
| 125 | + | let (workspace, rest) = rest.split_once('/')?; | |
| 126 | + | let workspace = workspace.to_ascii_lowercase(); | |
| 127 | + | if workspace.is_empty() { | |
| 128 | + | return None; | |
| 129 | + | } | |
| 130 | + | let parts: Vec<&str> = rest.split('/').collect(); | |
| 131 | + | let last = *parts.last()?; | |
| 132 | + | let (base, checksum) = split_checksum(last); | |
| 133 | + | let group_of = |parts: &[&str]| -> Option<String> { (!parts.is_empty() && parts.iter().all(|p| valid_group_part(p))).then(|| parts.join(".")) }; | |
| 134 | + | let n = parts.len(); | |
| 135 | + | if base == METADATA { | |
| 136 | + | if n >= 4 && is_snapshot(parts[n - 2]) && valid_version(parts[n - 2]) && valid_artifact(parts[n - 3]) | |
| 137 | + | && let Some(group) = group_of(&parts[..n - 3]) { | |
| 138 | + | let (artifact, version) = (parts[n - 3].to_owned(), parts[n - 2].to_owned()); | |
| 139 | + | return Some((workspace, MavenPath::VersionMetadata { group, artifact, version, checksum })); | |
| 140 | + | } | |
| 141 | + | if n < 3 || !valid_artifact(parts[n - 2]) { | |
| 142 | + | return None; | |
| 143 | + | } | |
| 144 | + | let group = group_of(&parts[..n - 2])?; | |
| 145 | + | return Some((workspace, MavenPath::ArtifactMetadata { group, artifact: parts[n - 2].to_owned(), checksum })); | |
| 146 | + | } | |
| 147 | + | if n < 4 { | |
| 148 | + | return None; | |
| 149 | + | } | |
| 150 | + | let (artifact, version) = (parts[n - 3], parts[n - 2]); | |
| 151 | + | if !valid_artifact(artifact) || !valid_version(version) { | |
| 152 | + | return None; | |
| 153 | + | } | |
| 154 | + | parse_file(artifact, version, base)?; | |
| 155 | + | let group = group_of(&parts[..n - 3])?; | |
| 156 | + | Some(( | |
| 157 | + | workspace, | |
| 158 | + | MavenPath::File { group, artifact: artifact.to_owned(), version: version.to_owned(), file: base.to_owned(), checksum }, | |
| 159 | + | )) | |
| 160 | + | } | |
| 161 | + | ||
| 162 | + | /// What a file's name says: `web-1.0.0-sources.jar` is the `sources` | |
| 163 | + | /// classifier's `jar`; `web-1.0-20261006.120000-3.pom` is build 3 of the | |
| 164 | + | /// SNAPSHOT `1.0-SNAPSHOT`, made at that time (UTC). | |
| 165 | + | #[derive(Clone, Debug, PartialEq, Eq)] | |
| 166 | + | pub struct FileName { | |
| 167 | + | pub classifier: Option<String>, | |
| 168 | + | pub extension: String, | |
| 169 | + | /// For a SNAPSHOT's timestamped file: `20261006.120000` and the build. | |
| 170 | + | pub build: Option<(String, u32)>, | |
| 171 | + | } | |
| 172 | + | ||
| 173 | + | impl FileName { | |
| 174 | + | /// The version its name holds: the SNAPSHOT's timestamped one. | |
| 175 | + | pub fn value(&self, version: &str) -> String { | |
| 176 | + | match &self.build { | |
| 177 | + | Some((timestamp, build)) => format!("{}-{timestamp}-{build}", version.trim_end_matches("-SNAPSHOT")), | |
| 178 | + | None => version.to_owned(), | |
| 179 | + | } | |
| 180 | + | } | |
| 181 | + | } | |
| 182 | + | ||
| 183 | + | /// `20261006.120000-3`: a SNAPSHOT build's time and number. | |
| 184 | + | fn snapshot_build(text: &str) -> Option<(String, u32, usize)> { | |
| 185 | + | let bytes = text.as_bytes(); | |
| 186 | + | if bytes.len() < 17 || !bytes[..8].iter().all(u8::is_ascii_digit) || bytes[8] != b'.' || !bytes[9..15].iter().all(u8::is_ascii_digit) || bytes[15] != b'-' { | |
| 187 | + | return None; | |
| 188 | + | } | |
| 189 | + | let digits = bytes[16..].iter().take_while(|b| b.is_ascii_digit()).count(); | |
| 190 | + | let build = text[16..16 + digits].parse().ok()?; | |
| 191 | + | Some((text[..15].to_owned(), build, 16 + digits)) | |
| 192 | + | } | |
| 193 | + | ||
| 194 | + | pub fn parse_file(artifact: &str, version: &str, file: &str) -> Option<FileName> { | |
| 195 | + | let rest = file.strip_prefix(artifact)?.strip_prefix('-')?; | |
| 196 | + | let (after, build) = if let Some(after) = rest.strip_prefix(version) { | |
| 197 | + | (after, None) | |
| 198 | + | } else { | |
| 199 | + | let base = version.strip_suffix("-SNAPSHOT")?; | |
| 200 | + | let stamped = rest.strip_prefix(base)?.strip_prefix('-')?; | |
| 201 | + | let (timestamp, build, used) = snapshot_build(stamped)?; | |
| 202 | + | (&stamped[used..], Some((timestamp, build))) | |
| 203 | + | }; | |
| 204 | + | let (classifier, extension) = if let Some(ext) = after.strip_prefix('.') { | |
| 205 | + | (None, ext) | |
| 206 | + | } else { | |
| 207 | + | let (classifier, ext) = after.strip_prefix('-')?.split_once('.')?; | |
| 208 | + | if classifier.is_empty() { | |
| 209 | + | return None; | |
| 210 | + | } | |
| 211 | + | (Some(classifier.to_owned()), ext) | |
| 212 | + | }; | |
| 213 | + | if extension.is_empty() || extension.contains('/') { | |
| 214 | + | return None; | |
| 215 | + | } | |
| 216 | + | Some(FileName { classifier, extension: extension.to_owned(), build }) | |
| 217 | + | } | |
| 218 | + | ||
| 219 | + | /// The media type a file is served with. | |
| 220 | + | pub fn media_type(file: &str) -> &'static str { | |
| 221 | + | let ext = file.rsplit('.').next().unwrap_or(""); | |
| 222 | + | match ext { | |
| 223 | + | "pom" | "xml" => "application/xml", | |
| 224 | + | "jar" | "war" | "ear" | "aar" => "application/java-archive", | |
| 225 | + | "module" | "json" => "application/json", | |
| 226 | + | "asc" => "application/pgp-signature", | |
| 227 | + | "zip" => "application/zip", | |
| 228 | + | _ => "application/octet-stream", | |
| 229 | + | } | |
| 230 | + | } | |
| 231 | + | ||
| 232 | + | #[derive(Clone, Debug, PartialEq, Eq, PartialOrd, Ord)] | |
| 233 | + | enum Item { | |
| 234 | + | /// A qualifier, by its rank, then by name for those Maven does not know. | |
| 235 | + | Qualifier(u8, String), | |
| 236 | + | Number(u64), | |
| 237 | + | } | |
| 238 | + | ||
| 239 | + | /// Maven's order of the qualifiers it knows; a release (no qualifier) is | |
| 240 | + | /// `ga`, `final` and `release`. | |
| 241 | + | fn rank(word: &str) -> (u8, String) { | |
| 242 | + | match word { | |
| 243 | + | "alpha" | "a" => (1, String::new()), | |
| 244 | + | "beta" | "b" => (2, String::new()), | |
| 245 | + | "milestone" | "m" => (3, String::new()), | |
| 246 | + | "rc" | "cr" => (4, String::new()), | |
| 247 | + | "snapshot" => (5, String::new()), | |
| 248 | + | "" | "ga" | "final" | "release" => (6, String::new()), | |
| 249 | + | "sp" => (7, String::new()), | |
| 250 | + | other => (8, other.to_owned()), | |
| 251 | + | } | |
| 252 | + | } | |
| 253 | + | ||
| 254 | + | fn items(version: &str) -> Vec<Item> { | |
| 255 | + | let lower = version.to_ascii_lowercase(); | |
| 256 | + | let mut out = Vec::new(); | |
| 257 | + | let mut word = String::new(); | |
| 258 | + | let mut digits: Option<bool> = None; | |
| 259 | + | let flush = |word: &mut String, out: &mut Vec<Item>, number: bool| { | |
| 260 | + | if number { | |
| 261 | + | out.push(Item::Number(word.parse().unwrap_or(u64::MAX))); | |
| 262 | + | } else { | |
| 263 | + | let (r, name) = rank(word); | |
| 264 | + | out.push(Item::Qualifier(r, name)); | |
| 265 | + | } | |
| 266 | + | word.clear(); | |
| 267 | + | }; | |
| 268 | + | for c in lower.chars() { | |
| 269 | + | if c == '.' || c == '-' || c == '_' { | |
| 270 | + | if let Some(number) = digits.take() { | |
| 271 | + | flush(&mut word, &mut out, number); | |
| 272 | + | } | |
| 273 | + | continue; | |
| 274 | + | } | |
| 275 | + | let is_digit = c.is_ascii_digit(); | |
| 276 | + | if digits.is_some_and(|d| d != is_digit) { | |
| 277 | + | flush(&mut word, &mut out, digits.unwrap_or(false)); | |
| 278 | + | } | |
| 279 | + | digits = Some(is_digit); | |
| 280 | + | word.push(c); | |
| 281 | + | } | |
| 282 | + | if let Some(number) = digits { | |
| 283 | + | flush(&mut word, &mut out, number); | |
| 284 | + | } | |
| 285 | + | // Trailing zeros and releases say nothing: 1.0 is 1 is 1.0.0-ga. | |
| 286 | + | while matches!(out.last(), Some(Item::Number(0)) | Some(Item::Qualifier(6, _))) { | |
| 287 | + | out.pop(); | |
| 288 | + | } | |
| 289 | + | out | |
| 290 | + | } | |
| 291 | + | ||
| 292 | + | /// Maven's order of versions: `1.0-alpha` < `1.0-rc1` < `1.0-SNAPSHOT` < | |
| 293 | + | /// `1.0` < `1.0.1` < `1.1`. | |
| 294 | + | pub fn compare(a: &str, b: &str) -> Ordering { | |
| 295 | + | let (a, b) = (items(a), items(b)); | |
| 296 | + | let missing = |item: &Item| match item { | |
| 297 | + | Item::Number(n) => (*n).cmp(&0), | |
| 298 | + | Item::Qualifier(r, _) => r.cmp(&6), | |
| 299 | + | }; | |
| 300 | + | for i in 0..a.len().max(b.len()) { | |
| 301 | + | let order = match (a.get(i), b.get(i)) { | |
| 302 | + | (Some(x), Some(y)) => x.cmp(y), | |
| 303 | + | (Some(x), None) => missing(x), | |
| 304 | + | (None, Some(y)) => missing(y).reverse(), | |
| 305 | + | (None, None) => Ordering::Equal, | |
| 306 | + | }; | |
| 307 | + | if order != Ordering::Equal { | |
| 308 | + | return order; | |
| 309 | + | } | |
| 310 | + | } | |
| 311 | + | Ordering::Equal | |
| 312 | + | } | |
| 313 | + | ||
| 314 | + | /// `yyyyMMddHHmmss` from an RFC 3339 time, as Maven writes `lastUpdated`. | |
| 315 | + | pub fn last_updated(rfc3339: &str) -> String { | |
| 316 | + | rfc3339.chars().filter(char::is_ascii_digit).take(14).collect() | |
| 317 | + | } | |
| 318 | + | ||
| 319 | + | /// An artifact's `maven-metadata.xml`: its versions in Maven's order, the | |
| 320 | + | /// highest as `latest`, and the highest that is not a SNAPSHOT as `release`. | |
| 321 | + | pub fn artifact_metadata(group: &str, artifact: &str, versions: &[String], updated: &str) -> String { | |
| 322 | + | let mut sorted: Vec<&String> = versions.iter().collect(); | |
| 323 | + | sorted.sort_by(|a, b| compare(a, b)); | |
| 324 | + | sorted.dedup(); | |
| 325 | + | let latest = sorted.last(); | |
| 326 | + | let release = sorted.iter().rev().find(|v| !is_snapshot(v)); | |
| 327 | + | let mut xml = String::from("<?xml version=\"1.0\" encoding=\"UTF-8\"?>\n<metadata>\n"); | |
| 328 | + | xml.push_str(&format!(" <groupId>{}</groupId>\n <artifactId>{}</artifactId>\n <versioning>\n", xml::escape(group), xml::escape(artifact))); | |
| 329 | + | if let Some(latest) = latest { | |
| 330 | + | xml.push_str(&format!(" <latest>{}</latest>\n", xml::escape(latest))); | |
| 331 | + | } | |
| 332 | + | if let Some(release) = release { | |
| 333 | + | xml.push_str(&format!(" <release>{}</release>\n", xml::escape(release))); | |
| 334 | + | } | |
| 335 | + | xml.push_str(" <versions>\n"); | |
| 336 | + | for version in &sorted { | |
| 337 | + | xml.push_str(&format!(" <version>{}</version>\n", xml::escape(version))); | |
| 338 | + | } | |
| 339 | + | xml.push_str(&format!(" </versions>\n <lastUpdated>{}</lastUpdated>\n </versioning>\n</metadata>\n", last_updated(updated))); | |
| 340 | + | xml | |
| 341 | + | } | |
| 342 | + | ||
| 343 | + | /// The newest build among a SNAPSHOT's files: its timestamp and number. | |
| 344 | + | pub fn newest_build(artifact: &str, version: &str, files: &[String]) -> Option<(String, u32)> { | |
| 345 | + | files | |
| 346 | + | .iter() | |
| 347 | + | .filter_map(|file| parse_file(artifact, version, file)?.build) | |
| 348 | + | .max_by(|a, b| (a.1, &a.0).cmp(&(b.1, &b.0))) | |
| 349 | + | } | |
| 350 | + | ||
| 351 | + | /// A SNAPSHOT version's `maven-metadata.xml`, from its files' names: the | |
| 352 | + | /// newest build, and the newest file of each classifier and extension. | |
| 353 | + | pub fn snapshot_metadata(group: &str, artifact: &str, version: &str, files: &[String]) -> Option<String> { | |
| 354 | + | let mut newest: Vec<(FileName, String)> = Vec::new(); | |
| 355 | + | for file in files { | |
| 356 | + | let Some(name) = parse_file(artifact, version, file) else { continue }; | |
| 357 | + | let Some(build) = name.build.clone() else { continue }; | |
| 358 | + | let key = (name.classifier.clone(), name.extension.clone()); | |
| 359 | + | match newest.iter_mut().find(|(n, _)| (n.classifier.clone(), n.extension.clone()) == key) { | |
| 360 | + | Some(kept) => { | |
| 361 | + | let had = kept.0.build.clone().unwrap_or_default(); | |
| 362 | + | if (build.1, &build.0) > (had.1, &had.0) { | |
| 363 | + | *kept = (name, file.clone()); | |
| 364 | + | } | |
| 365 | + | } | |
| 366 | + | None => newest.push((name, file.clone())), | |
| 367 | + | } | |
| 368 | + | } | |
| 369 | + | let (timestamp, build) = newest.iter().filter_map(|(n, _)| n.build.clone()).max_by(|a, b| (a.1, &a.0).cmp(&(b.1, &b.0)))?; | |
| 370 | + | newest.sort_by(|a, b| (&a.0.extension, &a.0.classifier).cmp(&(&b.0.extension, &b.0.classifier))); | |
| 371 | + | let updated = |stamp: &str| stamp.replace('.', ""); | |
| 372 | + | let mut xml = String::from("<?xml version=\"1.0\" encoding=\"UTF-8\"?>\n<metadata modelVersion=\"1.1.0\">\n"); | |
| 373 | + | xml.push_str(&format!( | |
| 374 | + | " <groupId>{}</groupId>\n <artifactId>{}</artifactId>\n <version>{}</version>\n <versioning>\n", | |
| 375 | + | xml::escape(group), | |
| 376 | + | xml::escape(artifact), | |
| 377 | + | xml::escape(version) | |
| 378 | + | )); | |
| 379 | + | xml.push_str(&format!(" <snapshot>\n <timestamp>{timestamp}</timestamp>\n <buildNumber>{build}</buildNumber>\n </snapshot>\n")); | |
| 380 | + | xml.push_str(&format!(" <lastUpdated>{}</lastUpdated>\n <snapshotVersions>\n", updated(×tamp))); | |
| 381 | + | for (name, _) in &newest { | |
| 382 | + | xml.push_str(" <snapshotVersion>\n"); | |
| 383 | + | if let Some(classifier) = &name.classifier { | |
| 384 | + | xml.push_str(&format!(" <classifier>{}</classifier>\n", xml::escape(classifier))); | |
| 385 | + | } | |
| 386 | + | let stamp = name.build.as_ref().map(|b| b.0.clone()).unwrap_or_default(); | |
| 387 | + | xml.push_str(&format!( | |
| 388 | + | " <extension>{}</extension>\n <value>{}</value>\n <updated>{}</updated>\n </snapshotVersion>\n", | |
| 389 | + | xml::escape(&name.extension), | |
| 390 | + | xml::escape(&name.value(version)), | |
| 391 | + | updated(&stamp) | |
| 392 | + | )); | |
| 393 | + | } | |
| 394 | + | xml.push_str(" </snapshotVersions>\n </versioning>\n</metadata>\n"); | |
| 395 | + | Some(xml) | |
| 396 | + | } | |
| 397 | + | ||
| 398 | + | /// What is read from a POM. | |
| 399 | + | #[derive(Clone, Debug, Default, PartialEq, Eq)] | |
| 400 | + | pub struct Pom { | |
| 401 | + | pub group: String, | |
| 402 | + | pub artifact: String, | |
| 403 | + | pub version: String, | |
| 404 | + | pub name: Option<String>, | |
| 405 | + | pub description: Option<String>, | |
| 406 | + | /// `<scm><url>`, else `<url>`: where its source is. | |
| 407 | + | pub source: Option<String>, | |
| 408 | + | } | |
| 409 | + | ||
| 410 | + | /// Reads a POM, with the groupId and version a `<parent>` gives it. | |
| 411 | + | pub fn read_pom(bytes: &[u8]) -> Result<Pom, String> { | |
| 412 | + | let text = std::str::from_utf8(bytes).map_err(|_| "The POM is not UTF-8.".to_owned())?; | |
| 413 | + | let project = xml::parse(text).map_err(|problem| format!("The POM is not XML: {problem}"))?; | |
| 414 | + | if project.name != "project" { | |
| 415 | + | return Err("The POM has no <project>.".to_owned()); | |
| 416 | + | } | |
| 417 | + | let parent = project.child("parent"); | |
| 418 | + | let inherited = |key: &str| project.child_text(key).or_else(|| parent.and_then(|p| p.child_text(key))); | |
| 419 | + | Ok(Pom { | |
| 420 | + | group: inherited("groupId").ok_or("The POM names no groupId.")?, | |
| 421 | + | artifact: project.child_text("artifactId").ok_or("The POM names no artifactId.")?, | |
| 422 | + | version: inherited("version").ok_or("The POM names no version.")?, | |
| 423 | + | name: project.child_text("name"), | |
| 424 | + | description: project.child_text("description"), | |
| 425 | + | source: project.child("scm").and_then(|scm| scm.child_text("url")).or_else(|| project.child_text("url")), | |
| 426 | + | }) | |
| 427 | + | } | |
| 428 | + | ||
| 429 | + | #[cfg(test)] | |
| 430 | + | mod tests { | |
| 431 | + | use super::*; | |
| 432 | + | ||
| 433 | + | fn file(group: &str, artifact: &str, version: &str, file: &str, checksum: Option<Checksum>) -> Option<(String, MavenPath)> { | |
| 434 | + | Some(( | |
| 435 | + | "acme".to_owned(), | |
| 436 | + | MavenPath::File { group: group.into(), artifact: artifact.into(), version: version.into(), file: file.into(), checksum }, | |
| 437 | + | )) | |
| 438 | + | } | |
| 439 | + | ||
| 440 | + | #[test] | |
| 441 | + | fn paths_are_the_standard_layout() { | |
| 442 | + | assert_eq!(route("/-/maven/Acme/com/acme/web/1.0.0/web-1.0.0.jar"), file("com.acme", "web", "1.0.0", "web-1.0.0.jar", None)); | |
| 443 | + | assert_eq!( | |
| 444 | + | route("/-/maven/acme/com/acme/web/1.0.0/web-1.0.0.pom.sha1"), | |
| 445 | + | file("com.acme", "web", "1.0.0", "web-1.0.0.pom", Some(Checksum::Sha1)) | |
| 446 | + | ); | |
| 447 | + | assert_eq!( | |
| 448 | + | route("/-/maven/acme/io/g1t/core-lib/2.1/core-lib-2.1-sources.jar.asc"), | |
| 449 | + | file("io.g1t", "core-lib", "2.1", "core-lib-2.1-sources.jar.asc", None) | |
| 450 | + | ); | |
| 451 | + | assert_eq!( | |
| 452 | + | route("/-/maven/acme/com/acme/web/1.0-SNAPSHOT/web-1.0-20261006.120000-3.module.sha512"), | |
| 453 | + | file("com.acme", "web", "1.0-SNAPSHOT", "web-1.0-20261006.120000-3.module", Some(Checksum::Sha512)) | |
| 454 | + | ); | |
| 455 | + | assert_eq!( | |
| 456 | + | route("/-/maven/acme/com/acme/web/maven-metadata.xml"), | |
| 457 | + | Some(("acme".into(), MavenPath::ArtifactMetadata { group: "com.acme".into(), artifact: "web".into(), checksum: None })) | |
| 458 | + | ); | |
| 459 | + | assert_eq!( | |
| 460 | + | route("/-/maven/acme/com/acme/web/maven-metadata.xml.md5"), | |
| 461 | + | Some(("acme".into(), MavenPath::ArtifactMetadata { group: "com.acme".into(), artifact: "web".into(), checksum: Some(Checksum::Md5) })) | |
| 462 | + | ); | |
| 463 | + | assert_eq!( | |
| 464 | + | route("/-/maven/acme/com/acme/web/1.0-SNAPSHOT/maven-metadata.xml"), | |
| 465 | + | Some(( | |
| 466 | + | "acme".into(), | |
| 467 | + | MavenPath::VersionMetadata { group: "com.acme".into(), artifact: "web".into(), version: "1.0-SNAPSHOT".into(), checksum: None } | |
| 468 | + | )) | |
| 469 | + | ); | |
| 470 | + | // Not a file of the artifact, or not in the layout. | |
| 471 | + | assert_eq!(route("/-/maven/acme/com/acme/web/1.0.0/other-1.0.0.jar"), None); | |
| 472 | + | assert_eq!(route("/-/maven/acme/com/acme/web/1.0.0/web-1.0.1"), None); | |
| 473 | + | assert_eq!(route("/-/maven/acme/web/1.0.0/web-1.0.0.jar"), None, "no groupId"); | |
| 474 | + | assert_eq!(route("/-/maven/acme/com/../web/1.0.0/web-1.0.0.jar"), None); | |
| 475 | + | assert_eq!(route("/-/maven/acme/com/acme/web/1.0.0/"), None); | |
| 476 | + | assert_eq!(route("/-/maven/acme/maven-metadata.xml"), None); | |
| 477 | + | assert_eq!(route("/-/maven/"), None); | |
| 478 | + | } | |
| 479 | + | ||
| 480 | + | #[test] | |
| 481 | + | fn file_names_say_their_classifier_extension_and_build() { | |
| 482 | + | assert_eq!(parse_file("web", "1.0.0", "web-1.0.0.jar"), Some(FileName { classifier: None, extension: "jar".into(), build: None })); | |
| 483 | + | assert_eq!( | |
| 484 | + | parse_file("web", "1.0.0", "web-1.0.0-javadoc.jar"), | |
| 485 | + | Some(FileName { classifier: Some("javadoc".into()), extension: "jar".into(), build: None }) | |
| 486 | + | ); | |
| 487 | + | assert_eq!(parse_file("web", "1.0.0", "web-1.0.0.tar.gz").unwrap().extension, "tar.gz"); | |
| 488 | + | let stamped = parse_file("web", "1.0-SNAPSHOT", "web-1.0-20261006.120000-12-sources.jar").unwrap(); | |
| 489 | + | assert_eq!(stamped.build, Some(("20261006.120000".into(), 12))); | |
| 490 | + | assert_eq!(stamped.classifier.as_deref(), Some("sources")); | |
| 491 | + | assert_eq!(stamped.value("1.0-SNAPSHOT"), "1.0-20261006.120000-12"); | |
| 492 | + | assert_eq!(parse_file("web", "1.0-SNAPSHOT", "web-1.0-SNAPSHOT.jar").unwrap().build, None); | |
| 493 | + | assert_eq!(parse_file("web", "1.0.0", "web-1.0.0"), None); | |
| 494 | + | assert_eq!(parse_file("web", "1.0.0", "web-1.0.0-.jar"), None); | |
| 495 | + | assert_eq!(parse_file("web", "1.0", "web-1.0-20261006.120000-1.jar").unwrap().build, None, "timestamps are a SNAPSHOT's"); | |
| 496 | + | } | |
| 497 | + | ||
| 498 | + | #[test] | |
| 499 | + | fn versions_sort_as_maven_sorts_them() { | |
| 500 | + | let mut versions = vec!["1.10", "1.0", "1.0-SNAPSHOT", "1.0-rc1", "1.0-alpha", "1.0.1", "1.2", "1.0-beta-2", "2.0.0-M1", "1.0-sp1"]; | |
| 501 | + | versions.sort_by(|a, b| compare(a, b)); | |
| 502 | + | assert_eq!(versions, ["1.0-alpha", "1.0-beta-2", "1.0-rc1", "1.0-SNAPSHOT", "1.0", "1.0-sp1", "1.0.1", "1.2", "1.10", "2.0.0-M1"]); | |
| 503 | + | assert_eq!(compare("1.0", "1.0.0"), Ordering::Equal); | |
| 504 | + | assert_eq!(compare("1.0-ga", "1"), Ordering::Equal); | |
| 505 | + | assert!(valid_version("1.0.0") && valid_version("2024.1+build") && !valid_version("-1") && !valid_version("1 0")); | |
| 506 | + | } | |
| 507 | + | ||
| 508 | + | #[test] | |
| 509 | + | fn artifact_metadata_lists_versions_with_latest_and_release() { | |
| 510 | + | let xml = artifact_metadata("com.acme", "web", &["1.1.0".into(), "1.0.0".into(), "1.2.0-SNAPSHOT".into()], "2026-10-06T12:30:05.000Z"); | |
| 511 | + | let doc = xml::parse(&xml).unwrap(); | |
| 512 | + | let versioning = doc.child("versioning").unwrap(); | |
| 513 | + | assert_eq!(doc.child_text("groupId").as_deref(), Some("com.acme")); | |
| 514 | + | assert_eq!(versioning.child_text("latest").as_deref(), Some("1.2.0-SNAPSHOT")); | |
| 515 | + | assert_eq!(versioning.child_text("release").as_deref(), Some("1.1.0")); | |
| 516 | + | let listed: Vec<String> = versioning.child("versions").unwrap().children_named("version").map(|v| v.text.clone()).collect(); | |
| 517 | + | assert_eq!(listed, ["1.0.0", "1.1.0", "1.2.0-SNAPSHOT"]); | |
| 518 | + | assert_eq!(versioning.child_text("lastUpdated").as_deref(), Some("20261006123005")); | |
| 519 | + | } | |
| 520 | + | ||
| 521 | + | #[test] | |
| 522 | + | fn snapshot_metadata_names_the_newest_build_of_each_file() { | |
| 523 | + | let files: Vec<String> = [ | |
| 524 | + | "web-1.0-20261006.120000-1.jar", | |
| 525 | + | "web-1.0-20261006.120000-1.pom", | |
| 526 | + | "web-1.0-20261007.090000-2.jar", | |
| 527 | + | "web-1.0-20261007.090000-2.pom", | |
| 528 | + | "web-1.0-20261006.120000-1-sources.jar", | |
| 529 | + | ] | |
| 530 | + | .iter() | |
| 531 | + | .map(|s| s.to_string()) | |
| 532 | + | .collect(); | |
| 533 | + | let xml = snapshot_metadata("com.acme", "web", "1.0-SNAPSHOT", &files).unwrap(); | |
| 534 | + | let doc = xml::parse(&xml).unwrap(); | |
| 535 | + | assert_eq!(doc.child_text("version").as_deref(), Some("1.0-SNAPSHOT")); | |
| 536 | + | let versioning = doc.child("versioning").unwrap(); | |
| 537 | + | let snapshot = versioning.child("snapshot").unwrap(); | |
| 538 | + | assert_eq!(snapshot.child_text("timestamp").as_deref(), Some("20261007.090000")); | |
| 539 | + | assert_eq!(snapshot.child_text("buildNumber").as_deref(), Some("2")); | |
| 540 | + | let shown: Vec<(Option<String>, String, String)> = versioning | |
| 541 | + | .child("snapshotVersions") | |
| 542 | + | .unwrap() | |
| 543 | + | .children_named("snapshotVersion") | |
| 544 | + | .map(|v| (v.child_text("classifier"), v.child_text("extension").unwrap(), v.child_text("value").unwrap())) | |
| 545 | + | .collect(); | |
| 546 | + | assert_eq!( | |
| 547 | + | shown, | |
| 548 | + | [ | |
| 549 | + | (None, "jar".to_owned(), "1.0-20261007.090000-2".to_owned()), | |
| 550 | + | (Some("sources".to_owned()), "jar".to_owned(), "1.0-20261006.120000-1".to_owned()), | |
| 551 | + | (None, "pom".to_owned(), "1.0-20261007.090000-2".to_owned()), | |
| 552 | + | ] | |
| 553 | + | ); | |
| 554 | + | assert_eq!(snapshot_metadata("com.acme", "web", "1.0-SNAPSHOT", &[]), None); | |
| 555 | + | assert_eq!(newest_build("web", "1.0-SNAPSHOT", &files), Some(("20261007.090000".to_owned(), 2))); | |
| 556 | + | assert_eq!(newest_build("web", "1.0-SNAPSHOT", &["web-1.0-SNAPSHOT.jar".to_owned()]), None); | |
| 557 | + | } | |
| 558 | + | ||
| 559 | + | #[test] | |
| 560 | + | fn a_pom_says_its_coordinates_and_source() { | |
| 561 | + | let pom = read_pom( | |
| 562 | + | br#"<?xml version="1.0"?> | |
| 563 | + | <project xmlns="http://maven.apache.org/POM/4.0.0"> | |
| 564 | + | <modelVersion>4.0.0</modelVersion> | |
| 565 | + | <parent><groupId>com.acme</groupId><artifactId>parent</artifactId><version>1.0.0</version></parent> | |
| 566 | + | <artifactId>web</artifactId> | |
| 567 | + | <name>Acme web</name> | |
| 568 | + | <description>The web client</description> | |
| 569 | + | <url>https://acme.example</url> | |
| 570 | + | <scm><url>https://g1t.sh/acme/web</url></scm> | |
| 571 | + | </project>"#, | |
| 572 | + | ) | |
| 573 | + | .unwrap(); | |
| 574 | + | assert_eq!((pom.group.as_str(), pom.artifact.as_str(), pom.version.as_str()), ("com.acme", "web", "1.0.0")); | |
| 575 | + | assert_eq!(pom.description.as_deref(), Some("The web client")); | |
| 576 | + | assert_eq!(pom.source.as_deref(), Some("https://g1t.sh/acme/web")); | |
| 577 | + | assert!(read_pom(b"<project><artifactId>x</artifactId></project>").is_err(), "no groupId"); | |
| 578 | + | assert!(read_pom(b"not xml").is_err()); | |
| 579 | + | } | |
| 580 | + | ||
| 581 | + | #[test] | |
| 582 | + | fn checksums_are_hex_of_the_file() { | |
| 583 | + | assert_eq!(split_checksum("web-1.0.jar.sha1"), ("web-1.0.jar", Some(Checksum::Sha1))); | |
| 584 | + | assert_eq!(split_checksum("web-1.0.jar"), ("web-1.0.jar", None)); | |
| 585 | + | assert_eq!(Checksum::Md5.of(b"abc"), "900150983cd24fb0d6963f7d28e17f72"); | |
| 586 | + | assert_eq!(Checksum::Sha1.of(b"abc"), "a9993e364706816aba3e25717850c26c9cd0d89d"); | |
| 587 | + | assert_eq!(sent_checksum(b"A9993E364706816ABA3E25717850C26C9CD0D89D web-1.0.jar\n"), "a9993e364706816aba3e25717850c26c9cd0d89d"); | |
| 588 | + | let sums = crate::db::Checksums::of(b"abc"); | |
| 589 | + | assert_eq!(Checksum::Sha1.pick(&sums, "x"), Checksum::Sha1.of(b"abc")); | |
| 590 | + | assert_eq!(Checksum::Sha512.pick(&sums, "x"), Checksum::Sha512.of(b"abc")); | |
| 591 | + | } | |
| 592 | + | } |
| 1 | + | //! The Maven repository: `g1t.sh/-/maven/<workspace>/`, one for each | |
| 2 | + | //! workspace, in the standard layout. `mvn deploy` and Gradle's `publish` | |
| 3 | + | //! upload each file with a `PUT` and Basic credentials (any username, a g1t | |
| 4 | + | //! token as the password); a `Bearer` token works too, for Gradle's header | |
| 5 | + | //! credentials. | |
| 6 | + | //! | |
| 7 | + | //! Files are kept once, by their SHA-256, with their MD5, SHA-1 and SHA-512 | |
| 8 | + | //! worked out as they arrive, so the checksum files beside them are | |
| 9 | + | //! answered without reading them again; checksums uploaded are checked | |
| 10 | + | //! against them, not kept. `maven-metadata.xml` is made from the versions | |
| 11 | + | //! on each read: an uploaded one is taken and let go. | |
| 12 | + | //! | |
| 13 | + | //! A release's files are written once. A SNAPSHOT's builds arrive as | |
| 14 | + | //! timestamped files beside each other, and its metadata names the newest. | |
| 15 | + | //! The POM is the version's record. The artifact's `maven-metadata.xml`, which | |
| 16 | + | //! Maven and Gradle upload last, publishes what the deploy brought, as an | |
| 17 | + | //! event and an audit entry. | |
| 18 | + | ||
| 19 | + | use g1t_contracts::User; | |
| 20 | + | use g1t_contracts::audit::AuditActor; | |
| 21 | + | use g1t_contracts::events::PackageEvent; | |
| 22 | + | use g1t_contracts::new_id; | |
| 23 | + | use g1t_kit::now_ms; | |
| 24 | + | use serde_json::{Value, json}; | |
| 25 | + | use worker::{Context, Headers, Method, Request, Response, ResponseBody, Result}; | |
| 26 | + | ||
| 27 | + | use crate::access::{self, Action}; | |
| 28 | + | use crate::db::{Checksums, NewFile, NewVersion, PackageRow, VersionRow}; | |
| 29 | + | use crate::digest::Digest; | |
| 30 | + | use crate::maven::{self, Checksum, MavenPath}; | |
| 31 | + | use crate::npm; | |
| 32 | + | use crate::oci::{Credentials, published_by}; | |
| 33 | + | use crate::store::BlobStore; | |
| 34 | + | use crate::{Caller, Packages, TargetOf}; | |
| 35 | + | ||
| 36 | + | const MAVEN: &str = "maven"; | |
| 37 | + | /// The most versions an artifact's metadata lists. | |
| 38 | + | const MAX_VERSIONS: u32 = 5000; | |
| 39 | + | /// The longest POM read for its description and source. | |
| 40 | + | const MAX_POM_BYTES: usize = 1024 * 1024; | |
| 41 | + | const DOCS: &str = "https://docs.g1t.sh/guides/maven/"; | |
| 42 | + | const TOKENS: &str = "https://g1t.sh/settings/tokens"; | |
| 43 | + | ||
| 44 | + | /// A plain-text answer. Maven and Gradle print the status; the text says | |
| 45 | + | /// why for anyone reading the response. | |
| 46 | + | fn error(status: u16, message: impl Into<String>) -> Result<Response> { | |
| 47 | + | let mut response = Response::ok(format!("{}\n", message.into()))?.with_status(status); | |
| 48 | + | response.headers_mut().set("content-type", "text/plain; charset=utf-8")?; | |
| 49 | + | if status == 401 { | |
| 50 | + | response.headers_mut().set("www-authenticate", "Basic realm=\"g1t\"")?; | |
| 51 | + | } | |
| 52 | + | Ok(response) | |
| 53 | + | } | |
| 54 | + | ||
| 55 | + | fn created() -> Result<Response> { | |
| 56 | + | Ok(Response::empty()?.with_status(201)) | |
| 57 | + | } | |
| 58 | + | ||
| 59 | + | fn sign_in() -> String { | |
| 60 | + | format!("Sign in to use this repository: Basic credentials with a g1t access token from {TOKENS} as the password. See {DOCS}") | |
| 61 | + | } | |
| 62 | + | ||
| 63 | + | /// A file or document as Maven reads it; a `HEAD` gets its headers alone. | |
| 64 | + | fn serve(bytes: Vec<u8>, media_type: &str, head: bool, cache: &str) -> Result<Response> { | |
| 65 | + | let headers = Headers::new(); | |
| 66 | + | headers.set("content-type", media_type)?; | |
| 67 | + | headers.set("content-length", &bytes.len().to_string())?; | |
| 68 | + | headers.set("cache-control", cache)?; | |
| 69 | + | let body = if head { ResponseBody::Empty } else { ResponseBody::Body(bytes) }; | |
| 70 | + | Ok(Response::from_body(body)?.with_headers(headers)) | |
| 71 | + | } | |
| 72 | + | ||
| 73 | + | impl Packages { | |
| 74 | + | /// Answers a Maven request. | |
| 75 | + | pub async fn maven(&self, request: Request, ctx: &Context) -> Result<Response> { | |
| 76 | + | let url = request.url()?; | |
| 77 | + | let Some((workspace, path)) = maven::route(url.path()) else { | |
| 78 | + | return error(404, "There is nothing at this address."); | |
| 79 | + | }; | |
| 80 | + | match self.maven_route(request, &workspace, path, ctx).await { | |
| 81 | + | Ok(response) => Ok(response), | |
| 82 | + | Err(problem) => { | |
| 83 | + | worker::console_error!("packages: maven {}: {problem}", url.path()); | |
| 84 | + | error(500, "Something went wrong on our side. Try again in a moment.") | |
| 85 | + | } | |
| 86 | + | } | |
| 87 | + | } | |
| 88 | + | ||
| 89 | + | async fn maven_route(&self, mut request: Request, workspace: &str, path: MavenPath, ctx: &Context) -> Result<Response> { | |
| 90 | + | let method = request.method(); | |
| 91 | + | let credentials = self.credentials(&request).await?; | |
| 92 | + | let read = matches!(method, Method::Get | Method::Head); | |
| 93 | + | if read && let Some(refused) = self.limited(&request, &credentials, "Basic credentials with a g1t token").await? { | |
| 94 | + | return Ok(refused); | |
| 95 | + | } | |
| 96 | + | let viewer = match credentials { | |
| 97 | + | Credentials::Viewer(viewer) => viewer, | |
| 98 | + | Credentials::None => None, | |
| 99 | + | Credentials::Token(_) | Credentials::Bad => { | |
| 100 | + | return error(401, format!("The token is not right, or has expired. Make an access token at {TOKENS}.")); | |
| 101 | + | } | |
| 102 | + | }; | |
| 103 | + | let viewer = viewer.as_ref(); | |
| 104 | + | let head = method == Method::Head; | |
| 105 | + | match (path, method) { | |
| 106 | + | (MavenPath::ArtifactMetadata { group, artifact, checksum }, Method::Get | Method::Head) => { | |
| 107 | + | self.maven_metadata(workspace, &group, &artifact, None, checksum, viewer, head).await | |
| 108 | + | } | |
| 109 | + | (MavenPath::VersionMetadata { group, artifact, version, checksum }, Method::Get | Method::Head) => { | |
| 110 | + | self.maven_metadata(workspace, &group, &artifact, Some(&version), checksum, viewer, head).await | |
| 111 | + | } | |
| 112 | + | (MavenPath::File { group, artifact, version, file, checksum }, Method::Get | Method::Head) => { | |
| 113 | + | self.maven_file(workspace, &maven::package_name(&group, &artifact), &version, &file, checksum, viewer, head, ctx) | |
| 114 | + | .await | |
| 115 | + | } | |
| 116 | + | (MavenPath::File { group, artifact, version, file, checksum: None }, Method::Put) => { | |
| 117 | + | self.maven_upload(&mut request, workspace, &group, &artifact, &version, &file, viewer).await | |
| 118 | + | } | |
| 119 | + | (MavenPath::File { group, artifact, version, file, checksum: Some(checksum) }, Method::Put) => { | |
| 120 | + | let name = maven::package_name(&group, &artifact); | |
| 121 | + | self.maven_checksum(&mut request, workspace, &name, &version, &file, checksum, viewer).await | |
| 122 | + | } | |
| 123 | + | (path @ (MavenPath::ArtifactMetadata { .. } | MavenPath::VersionMetadata { .. }), Method::Put) => { | |
| 124 | + | self.maven_metadata_upload(&mut request, workspace, &path, viewer).await | |
| 125 | + | } | |
| 126 | + | _ => error(405, "Not a method this address takes. Versions are deleted on the package's page."), | |
| 127 | + | } | |
| 128 | + | } | |
| 129 | + | ||
| 130 | + | async fn maven_package(&self, workspace: &str, name: &str) -> Result<Option<PackageRow>> { | |
| 131 | + | Ok(self.db.package(workspace, MAVEN, name).await?.filter(|p| !p.hidden())) | |
| 132 | + | } | |
| 133 | + | ||
| 134 | + | /// The answer for something that is not there: to someone not signed | |
| 135 | + | /// in, a 401 when the workspace has private artifacts (so Maven sends | |
| 136 | + | /// its credentials and asks again, and a private artifact looks like a | |
| 137 | + | /// missing one), else a 404. | |
| 138 | + | async fn maven_absent(&self, workspace: &str, viewer: Option<&User>) -> Result<Response> { | |
| 139 | + | if viewer.is_none() && self.db.has_private(workspace, MAVEN).await? { | |
| 140 | + | return error(401, sign_in()); | |
| 141 | + | } | |
| 142 | + | error(404, "Not found: no such artifact or file, or you cannot see it.") | |
| 143 | + | } | |
| 144 | + | ||
| 145 | + | /// Whether `viewer` may `action` the artifact, as the answer when not. | |
| 146 | + | async fn maven_check(&self, viewer: Option<&User>, package: &PackageRow, action: Action) -> Result<Option<Response>> { | |
| 147 | + | let target = TargetOf::package(package); | |
| 148 | + | let decision = access::decide(viewer, &target.view(), action); | |
| 149 | + | if decision.allowed { | |
| 150 | + | return Ok(None); | |
| 151 | + | } | |
| 152 | + | let readable = action != Action::Pull && access::decide(viewer, &target.view(), Action::Pull).allowed; | |
| 153 | + | if !readable && viewer.is_none() { | |
| 154 | + | return Ok(Some(error(401, sign_in())?)); | |
| 155 | + | } | |
| 156 | + | if !readable { | |
| 157 | + | return Ok(Some(self.maven_absent(&package.workspace, viewer).await?)); | |
| 158 | + | } | |
| 159 | + | Ok(Some(error(403, decision.reason.unwrap_or_else(|| "Not allowed.".to_owned()))?)) | |
| 160 | + | } | |
| 161 | + | ||
| 162 | + | /// `maven-metadata.xml` of an artifact, or of one of its SNAPSHOTs, or | |
| 163 | + | /// a checksum of it. | |
| 164 | + | #[allow(clippy::too_many_arguments)] | |
| 165 | + | async fn maven_metadata( | |
| 166 | + | &self, | |
| 167 | + | workspace: &str, | |
| 168 | + | group: &str, | |
| 169 | + | artifact: &str, | |
| 170 | + | snapshot: Option<&str>, | |
| 171 | + | checksum: Option<Checksum>, | |
| 172 | + | viewer: Option<&User>, | |
| 173 | + | head: bool, | |
| 174 | + | ) -> Result<Response> { | |
| 175 | + | let Some(package) = self.maven_package(workspace, &maven::package_name(group, artifact)).await? else { | |
| 176 | + | return self.maven_absent(workspace, viewer).await; | |
| 177 | + | }; | |
| 178 | + | if let Some(refusal) = self.maven_check(viewer, &package, Action::Pull).await? { | |
| 179 | + | return Ok(refusal); | |
| 180 | + | } | |
| 181 | + | let xml = match snapshot { | |
| 182 | + | None => { | |
| 183 | + | let versions: Vec<String> = self.db.versions(&package.id, MAX_VERSIONS).await?.into_iter().map(|v| v.version).collect(); | |
| 184 | + | if versions.is_empty() { | |
| 185 | + | return self.maven_absent(workspace, viewer).await; | |
| 186 | + | } | |
| 187 | + | maven::artifact_metadata(group, artifact, &versions, &package.updated_at) | |
| 188 | + | } | |
| 189 | + | Some(version) => { | |
| 190 | + | let Some(row) = self.db.version_named(&package.id, version).await? else { | |
| 191 | + | return self.maven_absent(workspace, viewer).await; | |
| 192 | + | }; | |
| 193 | + | let files: Vec<String> = self.db.files(&row.id).await?.into_iter().map(|f| f.name).collect(); | |
| 194 | + | let Some(xml) = maven::snapshot_metadata(group, artifact, version, &files) else { | |
| 195 | + | return self.maven_absent(workspace, viewer).await; | |
| 196 | + | }; | |
| 197 | + | xml | |
| 198 | + | } | |
| 199 | + | }; | |
| 200 | + | match checksum { | |
| 201 | + | Some(checksum) => serve(checksum.of(xml.as_bytes()).into_bytes(), "text/plain", head, "no-cache"), | |
| 202 | + | None => serve(xml.into_bytes(), "application/xml", head, "no-cache"), | |
| 203 | + | } | |
| 204 | + | } | |
| 205 | + | ||
| 206 | + | /// One of a version's files, or a checksum of it. | |
| 207 | + | #[allow(clippy::too_many_arguments)] | |
| 208 | + | async fn maven_file( | |
| 209 | + | &self, | |
| 210 | + | workspace: &str, | |
| 211 | + | name: &str, | |
| 212 | + | version: &str, | |
| 213 | + | file: &str, | |
| 214 | + | checksum: Option<Checksum>, | |
| 215 | + | viewer: Option<&User>, | |
| 216 | + | head: bool, | |
| 217 | + | ctx: &Context, | |
| 218 | + | ) -> Result<Response> { | |
| 219 | + | let Some(package) = self.maven_package(workspace, name).await? else { | |
| 220 | + | return self.maven_absent(workspace, viewer).await; | |
| 221 | + | }; | |
| 222 | + | if let Some(refusal) = self.maven_check(viewer, &package, Action::Pull).await? { | |
| 223 | + | return Ok(refusal); | |
| 224 | + | } | |
| 225 | + | let Some(row) = self.db.version_named(&package.id, version).await? else { | |
| 226 | + | return self.maven_absent(workspace, viewer).await; | |
| 227 | + | }; | |
| 228 | + | let Some(kept) = self.db.file(&row.id, file).await? else { | |
| 229 | + | return self.maven_absent(workspace, viewer).await; | |
| 230 | + | }; | |
| 231 | + | let Some(digest) = Digest::parse(&kept.digest) else { | |
| 232 | + | return self.maven_absent(workspace, viewer).await; | |
| 233 | + | }; | |
| 234 | + | let Some(blob) = self.db.package_blob(&package.id, &digest).await? else { | |
| 235 | + | return self.maven_absent(workspace, viewer).await; | |
| 236 | + | }; | |
| 237 | + | // A release's files never change; a SNAPSHOT's are timestamped, so | |
| 238 | + | // each name is one file too. | |
| 239 | + | let cache = "max-age=31536000"; | |
| 240 | + | if let Some(checksum) = checksum { | |
| 241 | + | let sums = match self.db.checksums(&digest).await? { | |
| 242 | + | Some(sums) => sums, | |
| 243 | + | None => { | |
| 244 | + | let Some(bytes) = self.store.read(&blob.object_key).await? else { | |
| 245 | + | return self.maven_absent(workspace, viewer).await; | |
| 246 | + | }; | |
| 247 | + | let sums = Checksums::of(&bytes); | |
| 248 | + | self.db.set_checksums(&digest, &sums).await?; | |
| 249 | + | sums | |
| 250 | + | } | |
| 251 | + | }; | |
| 252 | + | return serve(checksum.pick(&sums, digest.hex()).into_bytes(), "text/plain", head, cache); | |
| 253 | + | } | |
| 254 | + | let headers = Headers::new(); | |
| 255 | + | headers.set("content-type", maven::media_type(file))?; | |
| 256 | + | headers.set("content-length", &blob.size.to_string())?; | |
| 257 | + | headers.set("cache-control", cache)?; | |
| 258 | + | if head { | |
| 259 | + | return Ok(Response::from_body(ResponseBody::Empty)?.with_headers(headers)); | |
| 260 | + | } | |
| 261 | + | let Some(got) = self.store.get(&blob.object_key, None).await? else { | |
| 262 | + | return self.maven_absent(workspace, viewer).await; | |
| 263 | + | }; | |
| 264 | + | // A download is the artifact itself, not its POM, signature or | |
| 265 | + | // Gradle module file, which are read beside it. | |
| 266 | + | if let Some(parsed) = maven::parse_file(name.rsplit(':').next().unwrap_or(""), version, file) | |
| 267 | + | && parsed.classifier.is_none() | |
| 268 | + | && !matches!(parsed.extension.as_str(), "pom" | "module" | "asc") | |
| 269 | + | && !parsed.extension.ends_with(".asc") | |
| 270 | + | { | |
| 271 | + | self.count_download(&package.id, ctx); | |
| 272 | + | } | |
| 273 | + | Ok(Response::from_body(got.body)?.with_headers(headers)) | |
| 274 | + | } | |
| 275 | + | ||
| 276 | + | /// Who may upload to a new artifact: the repository its artifactId | |
| 277 | + | /// names, else the workspace. | |
| 278 | + | async fn maven_target(&self, workspace: &str, candidates: &[String]) -> Result<TargetOf> { | |
| 279 | + | let mut repo = None; | |
| 280 | + | for candidate in candidates { | |
| 281 | + | if let Some(found) = self.repo_by_name(workspace, candidate).await? { | |
| 282 | + | repo = Some(found); | |
| 283 | + | break; | |
| 284 | + | } | |
| 285 | + | } | |
| 286 | + | Ok(TargetOf { workspace: workspace.to_owned(), repo: repo.map(|r| (r.id, r.name, r.is_private)), public: false }) | |
| 287 | + | } | |
| 288 | + | ||
| 289 | + | /// Whether `viewer` may upload to the artifact (made on its first | |
| 290 | + | /// file), as the answer when not. | |
| 291 | + | fn maven_refusal(&self, viewer: Option<&User>, target: &TargetOf, exists: bool) -> Option<Result<Response>> { | |
| 292 | + | let decision = access::decide(viewer, &target.view(), Action::Push); | |
| 293 | + | if decision.allowed { | |
| 294 | + | return None; | |
| 295 | + | } | |
| 296 | + | if viewer.is_none() { | |
| 297 | + | return Some(error(401, sign_in())); | |
| 298 | + | } | |
| 299 | + | if exists && !access::decide(viewer, &target.view(), Action::Pull).allowed { | |
| 300 | + | return Some(error(404, "Not found: no such artifact, or you cannot see it.")); | |
| 301 | + | } | |
| 302 | + | Some(error(403, decision.reason.unwrap_or_else(|| "Not allowed.".to_owned()))) | |
| 303 | + | } | |
| 304 | + | ||
| 305 | + | async fn maven_body(&self, request: &mut Request) -> Result<std::result::Result<Vec<u8>, Response>> { | |
| 306 | + | let declared = request.headers().get("content-length")?.and_then(|n| n.parse::<u64>().ok()); | |
| 307 | + | let too_large = || { | |
| 308 | + | let mb = self.max_request / 1_000_000; | |
| 309 | + | error(413, format!("A file may be at most {mb} MB. See {DOCS}#size")) | |
| 310 | + | }; | |
| 311 | + | if declared.is_some_and(|n| n > self.max_request) { | |
| 312 | + | return Ok(Err(too_large()?)); | |
| 313 | + | } | |
| 314 | + | let bytes = request.bytes().await?; | |
| 315 | + | if bytes.len() as u64 > self.max_request { | |
| 316 | + | return Ok(Err(too_large()?)); | |
| 317 | + | } | |
| 318 | + | Ok(Ok(bytes)) | |
| 319 | + | } | |
| 320 | + | ||
| 321 | + | /// A `PUT` of one of a version's files. | |
| 322 | + | #[allow(clippy::too_many_arguments)] | |
| 323 | + | async fn maven_upload( | |
| 324 | + | &self, | |
| 325 | + | request: &mut Request, | |
| 326 | + | workspace: &str, | |
| 327 | + | group: &str, | |
| 328 | + | artifact: &str, | |
| 329 | + | version: &str, | |
| 330 | + | file: &str, | |
| 331 | + | viewer: Option<&User>, | |
| 332 | + | ) -> Result<Response> { | |
| 333 | + | let bytes = match self.maven_body(request).await? { | |
| 334 | + | Ok(bytes) => bytes, | |
| 335 | + | Err(refused) => return Ok(refused), | |
| 336 | + | }; | |
| 337 | + | if bytes.is_empty() { | |
| 338 | + | return error(400, format!("{file} is empty.")); | |
| 339 | + | } | |
| 340 | + | let Some(parsed) = maven::parse_file(artifact, version, file) else { | |
| 341 | + | return error(400, format!("{file} is not a file of {artifact} {version}.")); | |
| 342 | + | }; | |
| 343 | + | let is_pom = parsed.classifier.is_none() && parsed.extension == "pom"; | |
| 344 | + | let pom = if is_pom { | |
| 345 | + | let pom = match maven::read_pom(&bytes[..bytes.len().min(MAX_POM_BYTES)]) { | |
| 346 | + | Ok(pom) => pom, | |
| 347 | + | Err(message) => return error(400, message), | |
| 348 | + | }; | |
| 349 | + | if pom.group != group || pom.artifact != artifact || pom.version != version { | |
| 350 | + | return error( | |
| 351 | + | 400, | |
| 352 | + | format!( | |
| 353 | + | "The POM says {}:{}:{}, but it was uploaded as {group}:{artifact}:{version}.", | |
| 354 | + | pom.group, pom.artifact, pom.version | |
| 355 | + | ), | |
| 356 | + | ); | |
| 357 | + | } | |
| 358 | + | Some(pom) | |
| 359 | + | } else { | |
| 360 | + | None | |
| 361 | + | }; | |
| 362 | + | ||
| 363 | + | let name = maven::package_name(group, artifact); | |
| 364 | + | let found = self.db.package(workspace, MAVEN, &name).await?; | |
| 365 | + | if found.as_ref().is_some_and(PackageRow::hidden) || (found.is_none() && self.db.workspace_hidden(workspace).await?) { | |
| 366 | + | return error(403, format!("The workspace {workspace} is deleted; nothing can be published to it.")); | |
| 367 | + | } | |
| 368 | + | let target = match &found { | |
| 369 | + | Some(package) => TargetOf::package(package), | |
| 370 | + | None => { | |
| 371 | + | let lower = artifact.to_ascii_lowercase(); | |
| 372 | + | self.maven_target(workspace, &[lower]).await? | |
| 373 | + | } | |
| 374 | + | }; | |
| 375 | + | if let Some(refusal) = self.maven_refusal(viewer, &target, found.is_some()) { | |
| 376 | + | return refusal; | |
| 377 | + | } | |
| 378 | + | let caller = Caller { actor: viewer.map(AuditActor::of) }; | |
| 379 | + | let now = now_ms(); | |
| 380 | + | let package = match found { | |
| 381 | + | Some(package) => package, | |
| 382 | + | None => { | |
| 383 | + | self.db | |
| 384 | + | .create_package( | |
| 385 | + | &new_id("pkg", now), | |
| 386 | + | workspace, | |
| 387 | + | MAVEN, | |
| 388 | + | &name, | |
| 389 | + | target.repo.as_ref().map(|(id, repo, private)| (id.as_str(), repo.as_str(), *private)), | |
| 390 | + | caller.actor.as_ref().map_or("", |actor| actor.actor_id.as_str()), | |
| 391 | + | now, | |
| 392 | + | ) | |
| 393 | + | .await? | |
| 394 | + | } | |
| 395 | + | }; | |
| 396 | + | ||
| 397 | + | let digest = Digest::of(&bytes); | |
| 398 | + | let size = bytes.len() as u64; | |
| 399 | + | let (row, _) = self | |
| 400 | + | .db | |
| 401 | + | .version_or_new( | |
| 402 | + | &NewVersion { | |
| 403 | + | id: new_id("ver", now), | |
| 404 | + | package_id: package.id.clone(), | |
| 405 | + | version: version.to_owned(), | |
| 406 | + | digest: digest.to_string(), | |
| 407 | + | size: 0, | |
| 408 | + | metadata: "{}".to_owned(), | |
| 409 | + | subject: None, | |
| 410 | + | published_by: published_by(&caller), | |
| 411 | + | files: Vec::new(), | |
| 412 | + | }, | |
| 413 | + | now, | |
| 414 | + | ) | |
| 415 | + | .await?; | |
| 416 | + | if let Some(kept) = self.db.file(&row.id, file).await? { | |
| 417 | + | if kept.digest == digest.to_string() { | |
| 418 | + | return created(); | |
| 419 | + | } | |
| 420 | + | if !maven::is_snapshot(version) { | |
| 421 | + | return error( | |
| 422 | + | 409, | |
| 423 | + | format!("{file} is already published in {name} {version}, and a release's files are published once. Bump the version."), | |
| 424 | + | ); | |
| 425 | + | } | |
| 426 | + | } | |
| 427 | + | if let Some(refusal) = self.storage_refusal(&package, &[(digest.to_string(), size)]).await? { | |
| 428 | + | return error(403, refusal); | |
| 429 | + | } | |
| 430 | + | let stored = match self.db.blob(&digest).await? { | |
| 431 | + | Some(blob) => self.store.head(&blob.object_key).await?.is_some(), | |
| 432 | + | None => false, | |
| 433 | + | }; | |
| 434 | + | let sums = Checksums::of(&bytes); | |
| 435 | + | if !stored { | |
| 436 | + | self.store.put(&digest.object_key(), bytes).await?; | |
| 437 | + | } | |
| 438 | + | let media_type = maven::media_type(file); | |
| 439 | + | self.db.keep_blob(&package.id, &digest, size, Some(media_type), &digest.object_key(), now).await?; | |
| 440 | + | self.db.set_checksums(&digest, &sums).await?; | |
| 441 | + | self.db | |
| 442 | + | .put_file(&package.id, &row.id, &NewFile { name: file.to_owned(), digest: digest.to_string(), size, media_type: Some(media_type.to_owned()) }, now) | |
| 443 | + | .await?; | |
| 444 | + | self.db.measure(&package.workspace).await?; | |
| 445 | + | ||
| 446 | + | if let Some(pom) = pom { | |
| 447 | + | self.maven_pom(&package, &row, &digest, &pom, viewer).await?; | |
| 448 | + | } | |
| 449 | + | created() | |
| 450 | + | } | |
| 451 | + | ||
| 452 | + | /// The POM arrived: it is the version's record. Its description is | |
| 453 | + | /// the package's when it is the highest release, and a new artifact | |
| 454 | + | /// whose POM names its source on g1t is linked to that repository. | |
| 455 | + | async fn maven_pom(&self, package: &PackageRow, row: &VersionRow, digest: &Digest, pom: &maven::Pom, viewer: Option<&User>) -> Result<()> { | |
| 456 | + | let now = now_ms(); | |
| 457 | + | let version = row.version.as_str(); | |
| 458 | + | let mut metadata = row.meta(); | |
| 459 | + | if !metadata.is_object() { | |
| 460 | + | metadata = json!({}); | |
| 461 | + | } | |
| 462 | + | metadata["pom"] = json!(true); | |
| 463 | + | metadata["name"] = json!(pom.name); | |
| 464 | + | metadata["description"] = json!(pom.description); | |
| 465 | + | metadata["source"] = json!(pom.source); | |
| 466 | + | self.db.set_version(&row.id, &digest.to_string(), &metadata.to_string()).await?; | |
| 467 | + | let versions = self.db.versions(&package.id, MAX_VERSIONS).await?; | |
| 468 | + | let releases: Vec<&str> = versions.iter().map(|v| v.version.as_str()).filter(|v| !maven::is_snapshot(v)).collect(); | |
| 469 | + | let highest = if maven::is_snapshot(version) { | |
| 470 | + | releases.is_empty() | |
| 471 | + | } else { | |
| 472 | + | releases.iter().all(|other| maven::compare(other, version) != std::cmp::Ordering::Greater) | |
| 473 | + | }; | |
| 474 | + | if highest { | |
| 475 | + | self.db.set_readme(&package.id, None, pom.description.as_deref().or(pom.name.as_deref()), now).await?; | |
| 476 | + | } | |
| 477 | + | if package.repo_id.is_none() | |
| 478 | + | && versions.len() == 1 | |
| 479 | + | && let Some((owner, repo)) = pom.source.as_ref().and_then(|source| npm::repository_of(&Value::String(source.clone()), &self.host)) | |
| 480 | + | && owner == package.workspace | |
| 481 | + | && let Some(repo) = self.repo_by_name(&package.workspace, &repo).await? | |
| 482 | + | { | |
| 483 | + | let linked = TargetOf { | |
| 484 | + | workspace: package.workspace.clone(), | |
| 485 | + | repo: Some((repo.id.clone(), repo.name.clone(), repo.is_private)), | |
| 486 | + | public: false, | |
| 487 | + | }; | |
| 488 | + | if access::decide(viewer, &linked.view(), Action::Push).allowed { | |
| 489 | + | let visibility = if repo.is_private { "private" } else { "public" }; | |
| 490 | + | self.db.set_link(&package.id, Some((&repo.id, &repo.name)), visibility, now).await?; | |
| 491 | + | self.db.measure(&package.workspace).await?; | |
| 492 | + | } | |
| 493 | + | } | |
| 494 | + | Ok(()) | |
| 495 | + | } | |
| 496 | + | ||
| 497 | + | /// The artifact's `maven-metadata.xml` is what Maven and Gradle upload | |
| 498 | + | /// last: each version (or SNAPSHOT build) whose POM arrived since the | |
| 499 | + | /// last one is published now, with all its files, as an event and an | |
| 500 | + | /// audit entry. | |
| 501 | + | async fn maven_announce(&self, package: &PackageRow, caller: &Caller) -> Result<()> { | |
| 502 | + | let artifact = package.name.rsplit(':').next().unwrap_or("").to_owned(); | |
| 503 | + | for row in self.db.versions(&package.id, MAX_VERSIONS).await? { | |
| 504 | + | let mut metadata = row.meta(); | |
| 505 | + | if metadata["pom"] != json!(true) { | |
| 506 | + | continue; | |
| 507 | + | } | |
| 508 | + | let mark = if maven::is_snapshot(&row.version) { | |
| 509 | + | let files: Vec<String> = self.db.files(&row.id).await?.into_iter().map(|f| f.name).collect(); | |
| 510 | + | match maven::newest_build(&artifact, &row.version, &files) { | |
| 511 | + | Some((timestamp, build)) => format!("{timestamp}-{build}"), | |
| 512 | + | None => "unique".to_owned(), | |
| 513 | + | } | |
| 514 | + | } else { | |
| 515 | + | "release".to_owned() | |
| 516 | + | }; | |
| 517 | + | if metadata["published"].as_str() == Some(mark.as_str()) { | |
| 518 | + | continue; | |
| 519 | + | } | |
| 520 | + | metadata["published"] = json!(mark); | |
| 521 | + | self.db.set_version(&row.id, &row.digest, &metadata.to_string()).await?; | |
| 522 | + | let event = PackageEvent { | |
| 523 | + | version: Some(row.version.clone()), | |
| 524 | + | digest: Some(row.digest.clone()), | |
| 525 | + | size: Some(row.size), | |
| 526 | + | ..self.event_of(package) | |
| 527 | + | }; | |
| 528 | + | self.announce("package.published", package, event, caller).await; | |
| 529 | + | self.audit(caller, "package.publish", package, Some(&format!("{}/{}@{}", package.workspace, package.name, row.version)), None).await; | |
| 530 | + | } | |
| 531 | + | Ok(()) | |
| 532 | + | } | |
| 533 | + | ||
| 534 | + | /// A checksum uploaded beside a file: checked against the file's. | |
| 535 | + | #[allow(clippy::too_many_arguments)] | |
| 536 | + | async fn maven_checksum( | |
| 537 | + | &self, | |
| 538 | + | request: &mut Request, | |
| 539 | + | workspace: &str, | |
| 540 | + | name: &str, | |
| 541 | + | version: &str, | |
| 542 | + | file: &str, | |
| 543 | + | checksum: Checksum, | |
| 544 | + | viewer: Option<&User>, | |
| 545 | + | ) -> Result<Response> { | |
| 546 | + | let bytes = match self.maven_body(request).await? { | |
| 547 | + | Ok(bytes) => bytes, | |
| 548 | + | Err(refused) => return Ok(refused), | |
| 549 | + | }; | |
| 550 | + | let Some(package) = self.maven_package(workspace, name).await? else { | |
| 551 | + | return self.maven_absent(workspace, viewer).await; | |
| 552 | + | }; | |
| 553 | + | if let Some(refusal) = self.maven_check(viewer, &package, Action::Push).await? { | |
| 554 | + | return Ok(refusal); | |
| 555 | + | } | |
| 556 | + | let Some(row) = self.db.version_named(&package.id, version).await? else { | |
| 557 | + | return error(404, format!("Upload {file} before its checksum.")); | |
| 558 | + | }; | |
| 559 | + | let Some(kept) = self.db.file(&row.id, file).await? else { | |
| 560 | + | return error(404, format!("Upload {file} before its checksum.")); | |
| 561 | + | }; | |
| 562 | + | let Some(digest) = Digest::parse(&kept.digest) else { | |
| 563 | + | return error(404, format!("Upload {file} before its checksum.")); | |
| 564 | + | }; | |
| 565 | + | let Some(sums) = self.db.checksums(&digest).await? else { | |
| 566 | + | return created(); | |
| 567 | + | }; | |
| 568 | + | if maven::sent_checksum(&bytes) != checksum.pick(&sums, digest.hex()) { | |
| 569 | + | return error(400, format!("The {} checksum sent is not {file}'s. Upload the file again.", &checksum.suffix()[1..])); | |
| 570 | + | } | |
| 571 | + | created() | |
| 572 | + | } | |
| 573 | + | ||
| 574 | + | /// An uploaded `maven-metadata.xml`, or its checksum: the repository | |
| 575 | + | /// makes its own, so it is taken from anyone who may upload and let go. | |
| 576 | + | /// The artifact's own (not a SNAPSHOT's, nor its checksum) ends a | |
| 577 | + | /// deploy, and publishes what it brought. | |
| 578 | + | async fn maven_metadata_upload(&self, request: &mut Request, workspace: &str, path: &MavenPath, viewer: Option<&User>) -> Result<Response> { | |
| 579 | + | if let Err(refused) = self.maven_body(request).await? { | |
| 580 | + | return Ok(refused); | |
| 581 | + | } | |
| 582 | + | let (MavenPath::ArtifactMetadata { group, artifact, .. } | MavenPath::VersionMetadata { group, artifact, .. } | MavenPath::File { group, artifact, .. }) = path; | |
| 583 | + | let found = self.maven_package(workspace, &maven::package_name(group, artifact)).await?; | |
| 584 | + | let target = match &found { | |
| 585 | + | Some(package) => TargetOf::package(package), | |
| 586 | + | // A plugin group's metadata names no artifact of its own. | |
| 587 | + | None => self.maven_target(workspace, &[]).await?, | |
| 588 | + | }; | |
| 589 | + | if let Some(refusal) = self.maven_refusal(viewer, &target, found.is_some()) { | |
| 590 | + | return refusal; | |
| 591 | + | } | |
| 592 | + | if let (Some(package), MavenPath::ArtifactMetadata { checksum: None, .. }) = (&found, path) { | |
| 593 | + | let caller = Caller { actor: viewer.map(AuditActor::of) }; | |
| 594 | + | self.maven_announce(package, &caller).await?; | |
| 595 | + | } | |
| 596 | + | created() | |
| 597 | + | } | |
| 598 | + | } |
| 1 | + | //! What the NuGet feed needs that does not touch the network: package ids | |
| 2 | + | //! and NuGet's normalized versions, the feed's paths, the `.nuspec` read | |
| 3 | + | //! from a `.nupkg` (a zip), the multipart body `dotnet nuget push` sends, | |
| 4 | + | //! and the service index, registration and search documents of the v3 | |
| 5 | + | //! protocol. | |
| 6 | + | //! | |
| 7 | + | //! A version keeps what the documents need from its `.nuspec` as its | |
| 8 | + | //! metadata, made once when it is pushed. Unlisting (`dotnet nuget | |
| 9 | + | //! delete`) is the version's `yanked` column: an unlisted version is still | |
| 10 | + | //! downloaded by those who name it, but no longer searched or picked. | |
| 11 | + | ||
| 12 | + | use std::cmp::Ordering; | |
| 13 | + | ||
| 14 | + | use serde_json::{Value, json}; | |
| 15 | + | ||
| 16 | + | use crate::archive; | |
| 17 | + | use crate::xml; | |
| 18 | + | ||
| 19 | + | /// The longest id nuget.org takes. | |
| 20 | + | pub const MAX_ID: usize = 100; | |
| 21 | + | /// The largest `.nuspec` or README read from a package. | |
| 22 | + | const MAX_ENTRY_BYTES: usize = 1024 * 1024; | |
| 23 | + | ||
| 24 | + | /// An id: letters, digits and `_`, in parts joined by `.`, `-` or `_`. | |
| 25 | + | pub fn valid_id(id: &str) -> bool { | |
| 26 | + | let word = |b: u8| b.is_ascii_alphanumeric() || b == b'_'; | |
| 27 | + | let bytes = id.as_bytes(); | |
| 28 | + | !id.is_empty() | |
| 29 | + | && id.len() <= MAX_ID | |
| 30 | + | && word(bytes[0]) | |
| 31 | + | && word(bytes[bytes.len() - 1]) | |
| 32 | + | && bytes.iter().all(|b| word(*b) || matches!(b, b'.' | b'-')) | |
| 33 | + | && !bytes.windows(2).any(|w| matches!(w[0], b'.' | b'-') && matches!(w[1], b'.' | b'-')) | |
| 34 | + | } | |
| 35 | + | ||
| 36 | + | /// A version as NuGet reads it: up to four numbers, a pre-release label, | |
| 37 | + | /// and build metadata (which is not part of the version). | |
| 38 | + | #[derive(Clone, Debug, PartialEq, Eq)] | |
| 39 | + | struct Parsed { | |
| 40 | + | numbers: [u64; 4], | |
| 41 | + | release: Vec<String>, | |
| 42 | + | } | |
| 43 | + | ||
| 44 | + | fn parse(version: &str) -> Option<Parsed> { | |
| 45 | + | let version = version.trim(); | |
| 46 | + | let core = version.split_once('+').map_or(version, |(core, build)| { | |
| 47 | + | if build.is_empty() { "" } else { core } | |
| 48 | + | }); | |
| 49 | + | let (numbers, release) = core.split_once('-').map_or((core, None), |(n, r)| (n, Some(r))); | |
| 50 | + | let parts: Vec<&str> = numbers.split('.').collect(); | |
| 51 | + | if parts.is_empty() || parts.len() > 4 { | |
| 52 | + | return None; | |
| 53 | + | } | |
| 54 | + | let mut out = [0u64; 4]; | |
| 55 | + | for (i, part) in parts.iter().enumerate() { | |
| 56 | + | if part.is_empty() || !part.bytes().all(|b| b.is_ascii_digit()) || part.len() > 18 { | |
| 57 | + | return None; | |
| 58 | + | } | |
| 59 | + | out[i] = part.parse().ok()?; | |
| 60 | + | } | |
| 61 | + | let release = match release { | |
| 62 | + | None => Vec::new(), | |
| 63 | + | Some(text) => { | |
| 64 | + | let labels: Vec<String> = text.split('.').map(str::to_owned).collect(); | |
| 65 | + | if labels.iter().any(|l| l.is_empty() || !l.bytes().all(|b| b.is_ascii_alphanumeric() || b == b'-')) { | |
| 66 | + | return None; | |
| 67 | + | } | |
| 68 | + | labels | |
| 69 | + | } | |
| 70 | + | }; | |
| 71 | + | Some(Parsed { numbers: out, release }) | |
| 72 | + | } | |
| 73 | + | ||
| 74 | + | /// NuGet's normalized form: `1.0` is `1.0.0`, `1.0.0.0` is `1.0.0`, | |
| 75 | + | /// `01.2.3` is `1.2.3`, and `+build` is dropped. `None` for a version | |
| 76 | + | /// NuGet would not take. | |
| 77 | + | pub fn normalize(version: &str) -> Option<String> { | |
| 78 | + | let parsed = parse(version)?; | |
| 79 | + | let [a, b, c, d] = parsed.numbers; | |
| 80 | + | let mut out = if d == 0 { format!("{a}.{b}.{c}") } else { format!("{a}.{b}.{c}.{d}") }; | |
| 81 | + | if !parsed.release.is_empty() { | |
| 82 | + | out.push('-'); | |
| 83 | + | out.push_str(&parsed.release.join(".")); | |
| 84 | + | } | |
| 85 | + | Some(out) | |
| 86 | + | } | |
| 87 | + | ||
| 88 | + | pub fn is_prerelease(version: &str) -> bool { | |
| 89 | + | parse(version).is_some_and(|p| !p.release.is_empty()) | |
| 90 | + | } | |
| 91 | + | ||
| 92 | + | /// NuGet's order: by number, then a release above its pre-releases, whose | |
| 93 | + | /// labels compare as SemVer 2 says, ignoring case. | |
| 94 | + | pub fn compare(a: &str, b: &str) -> Ordering { | |
| 95 | + | let (Some(a), Some(b)) = (parse(a), parse(b)) else { | |
| 96 | + | return a.cmp(b); | |
| 97 | + | }; | |
| 98 | + | a.numbers.cmp(&b.numbers).then_with(|| match (a.release.is_empty(), b.release.is_empty()) { | |
| 99 | + | (true, true) => Ordering::Equal, | |
| 100 | + | (true, false) => Ordering::Greater, | |
| 101 | + | (false, true) => Ordering::Less, | |
| 102 | + | (false, false) => { | |
| 103 | + | for (x, y) in a.release.iter().zip(&b.release) { | |
| 104 | + | let order = match (x.parse::<u64>(), y.parse::<u64>()) { | |
| 105 | + | (Ok(x), Ok(y)) => x.cmp(&y), | |
| 106 | + | (Ok(_), Err(_)) => Ordering::Less, | |
| 107 | + | (Err(_), Ok(_)) => Ordering::Greater, | |
| 108 | + | (Err(_), Err(_)) => x.to_ascii_lowercase().cmp(&y.to_ascii_lowercase()), | |
| 109 | + | }; | |
| 110 | + | if order != Ordering::Equal { | |
| 111 | + | return order; | |
| 112 | + | } | |
| 113 | + | } | |
| 114 | + | a.release.len().cmp(&b.release.len()) | |
| 115 | + | } | |
| 116 | + | }) | |
| 117 | + | } | |
| 118 | + | ||
| 119 | + | /// Which of a version's files a flat container path asks for. | |
| 120 | + | #[derive(Clone, Copy, Debug, PartialEq, Eq)] | |
| 121 | + | pub enum Content { | |
| 122 | + | Nupkg, | |
| 123 | + | Nuspec, | |
| 124 | + | } | |
| 125 | + | ||
| 126 | + | /// One of the feed's endpoints, under `/-/nuget/<workspace>/`. | |
| 127 | + | #[derive(Clone, Debug, PartialEq, Eq)] | |
| 128 | + | pub enum NugetRoute { | |
| 129 | + | /// `v3/index.json`: the service index. | |
| 130 | + | Index, | |
| 131 | + | /// `v3/flatcontainer/<id>/index.json`: every version, lowercased. | |
| 132 | + | Versions { id: String }, | |
| 133 | + | /// `v3/flatcontainer/<id>/<version>/<id>.<version>.nupkg` or `<id>.nuspec`. | |
| 134 | + | Content { id: String, version: String, file: Content }, | |
| 135 | + | /// `v3/registration/<id>/index.json`. | |
| 136 | + | Registration { id: String }, | |
| 137 | + | /// `v3/registration/<id>/<version>.json`. | |
| 138 | + | Leaf { id: String, version: String }, | |
| 139 | + | /// `v3/query`. | |
| 140 | + | Search, | |
| 141 | + | /// `api/v2/package`: `dotnet nuget push`. | |
| 142 | + | Push, | |
| 143 | + | /// `api/v2/package/<id>/<version>`: `DELETE` unlists, `POST` lists again. | |
| 144 | + | Listing { id: String, version: String }, | |
| 145 | + | } | |
| 146 | + | ||
| 147 | + | /// The workspace and endpoint a path is. Ids are checked; versions are | |
| 148 | + | /// checked by the handler, which reads them as NuGet does. | |
| 149 | + | pub fn route(path: &str) -> Option<(String, NugetRoute)> { | |
| 150 | + | let rest = path.strip_prefix("/-/nuget/")?; | |
| 151 | + | let (workspace, rest) = rest.split_once('/')?; | |
| 152 | + | let workspace = workspace.to_ascii_lowercase(); | |
| 153 | + | if workspace.is_empty() { | |
| 154 | + | return None; | |
| 155 | + | } | |
| 156 | + | let parts: Vec<&str> = rest.trim_end_matches('/').split('/').collect(); | |
| 157 | + | let id = |text: &str| valid_id(text).then(|| text.to_owned()); | |
| 158 | + | let route = match parts.as_slice() { | |
| 159 | + | ["v3", "index.json"] => NugetRoute::Index, | |
| 160 | + | ["v3", "query"] => NugetRoute::Search, | |
| 161 | + | ["v3", "flatcontainer", name, "index.json"] => NugetRoute::Versions { id: id(name)? }, | |
| 162 | + | ["v3", "flatcontainer", name, version, file] => { | |
| 163 | + | let lower = format!("{}.{}", name.to_ascii_lowercase(), version.to_ascii_lowercase()); | |
| 164 | + | let file = if file.eq_ignore_ascii_case(&format!("{lower}.nupkg")) { | |
| 165 | + | Content::Nupkg | |
| 166 | + | } else if file.eq_ignore_ascii_case(&format!("{name}.nuspec")) { | |
| 167 | + | Content::Nuspec | |
| 168 | + | } else { | |
| 169 | + | return None; | |
| 170 | + | }; | |
| 171 | + | NugetRoute::Content { id: id(name)?, version: (*version).to_owned(), file } | |
| 172 | + | } | |
| 173 | + | ["v3", "registration", name, "index.json"] => NugetRoute::Registration { id: id(name)? }, | |
| 174 | + | ["v3", "registration", name, leaf] => NugetRoute::Leaf { id: id(name)?, version: leaf.strip_suffix(".json")?.to_owned() }, | |
| 175 | + | ["api", "v2", "package"] => NugetRoute::Push, | |
| 176 | + | ["api", "v2", "package", name, version] => NugetRoute::Listing { id: id(name)?, version: (*version).to_owned() }, | |
| 177 | + | _ => return None, | |
| 178 | + | }; | |
| 179 | + | Some((workspace, route)) | |
| 180 | + | } | |
| 181 | + | ||
| 182 | + | /// The `.nupkg` file in a `multipart/form-data` body, as `dotnet nuget | |
| 183 | + | /// push` sends it; a body that is not multipart is taken as the file. | |
| 184 | + | pub fn pushed_file<'a>(content_type: Option<&str>, body: &'a [u8]) -> Result<&'a [u8], String> { | |
| 185 | + | let Some(content_type) = content_type.filter(|c| c.to_ascii_lowercase().starts_with("multipart/")) else { | |
| 186 | + | return Ok(body); | |
| 187 | + | }; | |
| 188 | + | let boundary = content_type | |
| 189 | + | .split(';') | |
| 190 | + | .filter_map(|part| part.trim().split_once('=')) | |
| 191 | + | .find(|(key, _)| key.trim().eq_ignore_ascii_case("boundary")) | |
| 192 | + | .map(|(_, value)| value.trim().trim_matches('"')) | |
| 193 | + | .filter(|b| !b.is_empty()) | |
| 194 | + | .ok_or("The upload names no multipart boundary.")?; | |
| 195 | + | let find = |haystack: &[u8], needle: &[u8], from: usize| { | |
| 196 | + | haystack.get(from..).and_then(|rest| rest.windows(needle.len()).position(|w| w == needle)).map(|at| at + from) | |
| 197 | + | }; | |
| 198 | + | let delimiter = format!("--{boundary}"); | |
| 199 | + | let start = find(body, delimiter.as_bytes(), 0).ok_or("The upload holds no package.")?; | |
| 200 | + | let headers_end = find(body, b"\r\n\r\n", start).ok_or("The upload holds no package.")? + 4; | |
| 201 | + | let end = find(body, format!("\r\n{delimiter}").as_bytes(), headers_end).ok_or("The upload ends early.")?; | |
| 202 | + | Ok(&body[headers_end..end]) | |
| 203 | + | } | |
| 204 | + | ||
| 205 | + | /// A dependency group: the framework it is for (none for every one), and | |
| 206 | + | /// each dependency's id and version range. | |
| 207 | + | #[derive(Clone, Debug, PartialEq, Eq)] | |
| 208 | + | pub struct Group { | |
| 209 | + | pub target_framework: Option<String>, | |
| 210 | + | pub dependencies: Vec<(String, String)>, | |
| 211 | + | } | |
| 212 | + | ||
| 213 | + | /// What is read from a `.nuspec`. | |
| 214 | + | #[derive(Clone, Debug, Default, PartialEq, Eq)] | |
| 215 | + | pub struct Nuspec { | |
| 216 | + | pub id: String, | |
| 217 | + | pub version: String, | |
| 218 | + | pub title: Option<String>, | |
| 219 | + | pub description: Option<String>, | |
| 220 | + | pub summary: Option<String>, | |
| 221 | + | pub authors: Option<String>, | |
| 222 | + | pub tags: Vec<String>, | |
| 223 | + | pub project_url: Option<String>, | |
| 224 | + | pub repository_url: Option<String>, | |
| 225 | + | pub license_expression: Option<String>, | |
| 226 | + | pub license_url: Option<String>, | |
| 227 | + | pub icon_url: Option<String>, | |
| 228 | + | pub readme: Option<String>, | |
| 229 | + | pub require_license_acceptance: bool, | |
| 230 | + | pub groups: Vec<Group>, | |
| 231 | + | } | |
| 232 | + | ||
| 233 | + | /// A dependency's `version` as a range: `1.0` (at least 1.0) is | |
| 234 | + | /// `[1.0, )`; an interval is kept; none is any version. | |
| 235 | + | pub fn range(version: Option<&str>) -> String { | |
| 236 | + | match version.map(str::trim).filter(|v| !v.is_empty()) { | |
| 237 | + | None => "(, )".to_owned(), | |
| 238 | + | Some(v) if v.starts_with('[') || v.starts_with('(') => v.to_owned(), | |
| 239 | + | Some(v) => format!("[{v}, )"), | |
| 240 | + | } | |
| 241 | + | } | |
| 242 | + | ||
| 243 | + | pub fn read_nuspec(text: &str) -> Result<Nuspec, String> { | |
| 244 | + | let root = xml::parse(text).map_err(|problem| format!("The .nuspec is not XML: {problem}"))?; | |
| 245 | + | let metadata = root.child("metadata").ok_or("The .nuspec has no <metadata>.")?; | |
| 246 | + | let dependency = |d: &xml::Element| d.attribute("id").map(|id| (id.to_owned(), range(d.attribute("version")))); | |
| 247 | + | let mut groups = Vec::new(); | |
| 248 | + | if let Some(deps) = metadata.child("dependencies") { | |
| 249 | + | let loose: Vec<_> = deps.children_named("dependency").filter_map(dependency).collect(); | |
| 250 | + | if !loose.is_empty() { | |
| 251 | + | groups.push(Group { target_framework: None, dependencies: loose }); | |
| 252 | + | } | |
| 253 | + | for group in deps.children_named("group") { | |
| 254 | + | groups.push(Group { | |
| 255 | + | target_framework: group.attribute("targetFramework").map(str::to_owned).filter(|t| !t.is_empty()), | |
| 256 | + | dependencies: group.children_named("dependency").filter_map(dependency).collect(), | |
| 257 | + | }); | |
| 258 | + | } | |
| 259 | + | } | |
| 260 | + | let license = metadata.child("license"); | |
| 261 | + | Ok(Nuspec { | |
| 262 | + | id: metadata.child_text("id").ok_or("The .nuspec has no <id>.")?, | |
| 263 | + | version: metadata.child_text("version").ok_or("The .nuspec has no <version>.")?, | |
| 264 | + | title: metadata.child_text("title"), | |
| 265 | + | description: metadata.child_text("description"), | |
| 266 | + | summary: metadata.child_text("summary"), | |
| 267 | + | authors: metadata.child_text("authors"), | |
| 268 | + | tags: metadata.child_text("tags").map(|t| t.split([' ', ',', ';']).filter(|t| !t.is_empty()).map(str::to_owned).collect()).unwrap_or_default(), | |
| 269 | + | project_url: metadata.child_text("projectUrl"), | |
| 270 | + | repository_url: metadata.child("repository").and_then(|r| r.attribute("url")).map(str::to_owned).filter(|u| !u.is_empty()), | |
| 271 | + | license_expression: license | |
| 272 | + | .filter(|l| l.attribute("type") == Some("expression")) | |
| 273 | + | .map(|l| l.text.trim().to_owned()) | |
| 274 | + | .filter(|l| !l.is_empty()), | |
| 275 | + | license_url: metadata.child_text("licenseUrl"), | |
| 276 | + | icon_url: metadata.child_text("iconUrl"), | |
| 277 | + | readme: metadata.child_text("readme"), | |
| 278 | + | require_license_acceptance: metadata.child_text("requireLicenseAcceptance").is_some_and(|v| v.eq_ignore_ascii_case("true")), | |
| 279 | + | groups, | |
| 280 | + | }) | |
| 281 | + | } | |
| 282 | + | ||
| 283 | + | /// A package's `.nuspec` (read and as its bytes) and the README it names. | |
| 284 | + | pub struct Package { | |
| 285 | + | pub nuspec: Nuspec, | |
| 286 | + | pub nuspec_bytes: Vec<u8>, | |
| 287 | + | pub readme: Option<String>, | |
| 288 | + | } | |
| 289 | + | ||
| 290 | + | /// Reads a `.nupkg`: the `.nuspec` at its root, and its README. | |
| 291 | + | pub fn read_package(nupkg: &[u8]) -> Result<Package, String> { | |
| 292 | + | let entries = archive::zip_entries(nupkg).map_err(|_| "The package is not a .nupkg: it is not a zip.".to_owned())?; | |
| 293 | + | let entry = entries | |
| 294 | + | .iter() | |
| 295 | + | .find(|e| !e.name.contains('/') && e.name.to_ascii_lowercase().ends_with(".nuspec")) | |
| 296 | + | .ok_or("The package has no .nuspec.")?; | |
| 297 | + | let nuspec_bytes = archive::zip_read(nupkg, entry, MAX_ENTRY_BYTES)?; | |
| 298 | + | let text = String::from_utf8(nuspec_bytes.clone()).map_err(|_| "The .nuspec is not UTF-8.".to_owned())?; | |
| 299 | + | let nuspec = read_nuspec(&text)?; | |
| 300 | + | let readme = match &nuspec.readme { | |
| 301 | + | Some(path) => { | |
| 302 | + | let wanted = path.replace('\\', "/").trim_start_matches('/').to_ascii_lowercase(); | |
| 303 | + | entries | |
| 304 | + | .iter() | |
| 305 | + | .find(|e| e.name.to_ascii_lowercase() == wanted) | |
| 306 | + | .and_then(|e| archive::zip_read(nupkg, e, MAX_ENTRY_BYTES).ok()) | |
| 307 | + | .and_then(|bytes| String::from_utf8(bytes).ok()) | |
| 308 | + | } | |
| 309 | + | None => None, | |
| 310 | + | }; | |
| 311 | + | Ok(Package { nuspec, nuspec_bytes, readme }) | |
| 312 | + | } | |
| 313 | + | ||
| 314 | + | /// What a version keeps from its `.nuspec`, for the feed's documents. | |
| 315 | + | pub fn stored(nuspec: &Nuspec, version: &str) -> Value { | |
| 316 | + | json!({ | |
| 317 | + | "id": nuspec.id, | |
| 318 | + | "version": version, | |
| 319 | + | "title": nuspec.title, | |
| 320 | + | "description": nuspec.description, | |
| 321 | + | "summary": nuspec.summary, | |
| 322 | + | "authors": nuspec.authors, | |
| 323 | + | "tags": nuspec.tags, | |
| 324 | + | "project_url": nuspec.project_url, | |
| 325 | + | "license_expression": nuspec.license_expression, | |
| 326 | + | "license_url": nuspec.license_url, | |
| 327 | + | "icon_url": nuspec.icon_url, | |
| 328 | + | "require_license_acceptance": nuspec.require_license_acceptance, | |
| 329 | + | "dependency_groups": nuspec.groups.iter().map(|g| json!({ | |
| 330 | + | "target_framework": g.target_framework, | |
| 331 | + | "dependencies": g.dependencies.iter().map(|(id, range)| json!({ "id": id, "range": range })).collect::<Vec<_>>(), | |
| 332 | + | })).collect::<Vec<_>>(), | |
| 333 | + | }) | |
| 334 | + | } | |
| 335 | + | ||
| 336 | + | /// The service index: where the client finds each resource, under `base` | |
| 337 | + | /// (`https://g1t.sh/-/nuget/acme`). | |
| 338 | + | pub fn service_index(base: &str) -> Value { | |
| 339 | + | let resource = |id: String, kind: &str| json!({ "@id": id, "@type": kind }); | |
| 340 | + | let registration = format!("{base}/v3/registration/"); | |
| 341 | + | let query = format!("{base}/v3/query"); | |
| 342 | + | json!({ | |
| 343 | + | "version": "3.0.0", | |
| 344 | + | "resources": [ | |
| 345 | + | resource(format!("{base}/v3/flatcontainer/"), "PackageBaseAddress/3.0.0"), | |
| 346 | + | resource(registration.clone(), "RegistrationsBaseUrl"), | |
| 347 | + | resource(registration.clone(), "RegistrationsBaseUrl/3.0.0-rc"), | |
| 348 | + | resource(registration.clone(), "RegistrationsBaseUrl/3.0.0-beta"), | |
| 349 | + | resource(registration.clone(), "RegistrationsBaseUrl/3.4.0"), | |
| 350 | + | resource(registration, "RegistrationsBaseUrl/3.6.0"), | |
| 351 | + | resource(query.clone(), "SearchQueryService"), | |
| 352 | + | resource(query.clone(), "SearchQueryService/3.0.0-rc"), | |
| 353 | + | resource(query.clone(), "SearchQueryService/3.0.0-beta"), | |
| 354 | + | resource(query, "SearchQueryService/3.5.0"), | |
| 355 | + | resource(format!("{base}/api/v2/package"), "PackagePublish/2.0.0"), | |
| 356 | + | ], | |
| 357 | + | }) | |
| 358 | + | } | |
| 359 | + | ||
| 360 | + | /// One version as the registration and search documents list it. | |
| 361 | + | pub struct Listed<'a> { | |
| 362 | + | pub version: &'a str, | |
| 363 | + | pub metadata: &'a Value, | |
| 364 | + | pub published: &'a str, | |
| 365 | + | pub listed: bool, | |
| 366 | + | pub downloads: u64, | |
| 367 | + | } | |
| 368 | + | ||
| 369 | + | /// The addresses of a version's documents and files. | |
| 370 | + | pub struct Addresses { | |
| 371 | + | pub registration: String, | |
| 372 | + | pub leaf: String, | |
| 373 | + | pub content: String, | |
| 374 | + | } | |
| 375 | + | ||
| 376 | + | pub fn addresses(base: &str, id: &str, version: &str) -> Addresses { | |
| 377 | + | let (id, version) = (id.to_ascii_lowercase(), version.to_ascii_lowercase()); | |
| 378 | + | Addresses { | |
| 379 | + | registration: format!("{base}/v3/registration/{id}/index.json"), | |
| 380 | + | leaf: format!("{base}/v3/registration/{id}/{version}.json"), | |
| 381 | + | content: format!("{base}/v3/flatcontainer/{id}/{version}/{id}.{version}.nupkg"), | |
| 382 | + | } | |
| 383 | + | } | |
| 384 | + | ||
| 385 | + | /// A version's registration leaf, with its catalog entry inlined. | |
| 386 | + | pub fn leaf(base: &str, id: &str, listed: &Listed<'_>) -> Value { | |
| 387 | + | let at = addresses(base, id, listed.version); | |
| 388 | + | let m = listed.metadata; | |
| 389 | + | let groups: Vec<Value> = m["dependency_groups"] | |
| 390 | + | .as_array() | |
| 391 | + | .map(|groups| { | |
| 392 | + | groups | |
| 393 | + | .iter() | |
| 394 | + | .enumerate() | |
| 395 | + | .map(|(n, g)| { | |
| 396 | + | let deps: Vec<Value> = g["dependencies"] | |
| 397 | + | .as_array() | |
| 398 | + | .map(|deps| deps.iter().map(|d| json!({ "@id": format!("{}#dependency/{n}/{}", at.leaf, d["id"].as_str().unwrap_or("")), "id": d["id"], "range": d["range"] })).collect()) | |
| 399 | + | .unwrap_or_default(); | |
| 400 | + | let mut group = json!({ "@id": format!("{}#dependencygroup/{n}", at.leaf), "dependencies": deps }); | |
| 401 | + | if let Some(target) = g["target_framework"].as_str() { | |
| 402 | + | group["targetFramework"] = json!(target); | |
| 403 | + | } | |
| 404 | + | group | |
| 405 | + | }) | |
| 406 | + | .collect() | |
| 407 | + | }) | |
| 408 | + | .unwrap_or_default(); | |
| 409 | + | let text = |key: &str| m[key].as_str().unwrap_or("").to_owned(); | |
| 410 | + | json!({ | |
| 411 | + | "@id": at.leaf, | |
| 412 | + | "@type": "Package", | |
| 413 | + | "catalogEntry": { | |
| 414 | + | "@id": format!("{}#catalog", at.leaf), | |
| 415 | + | "@type": "PackageDetails", | |
| 416 | + | "id": m["id"].as_str().unwrap_or(id), | |
| 417 | + | "version": listed.version, | |
| 418 | + | "title": text("title"), | |
| 419 | + | "description": text("description"), | |
| 420 | + | "summary": text("summary"), | |
| 421 | + | "authors": text("authors"), | |
| 422 | + | "tags": m["tags"].as_array().cloned().unwrap_or_default(), | |
| 423 | + | "projectUrl": text("project_url"), | |
| 424 | + | "licenseExpression": text("license_expression"), | |
| 425 | + | "licenseUrl": text("license_url"), | |
| 426 | + | "iconUrl": text("icon_url"), | |
| 427 | + | "requireLicenseAcceptance": m["require_license_acceptance"].as_bool().unwrap_or(false), | |
| 428 | + | "dependencyGroups": groups, | |
| 429 | + | "listed": listed.listed, | |
| 430 | + | "published": listed.published, | |
| 431 | + | "packageContent": at.content, | |
| 432 | + | }, | |
| 433 | + | "packageContent": at.content, | |
| 434 | + | "registration": at.registration, | |
| 435 | + | }) | |
| 436 | + | } | |
| 437 | + | ||
| 438 | + | /// The registration index: every version, oldest first, in one page. | |
| 439 | + | pub fn registration(base: &str, id: &str, versions: &[Listed<'_>]) -> Value { | |
| 440 | + | let index = format!("{base}/v3/registration/{}/index.json", id.to_ascii_lowercase()); | |
| 441 | + | let (lower, upper) = (versions.first().map_or("", |v| v.version), versions.last().map_or("", |v| v.version)); | |
| 442 | + | json!({ | |
| 443 | + | "@id": index, | |
| 444 | + | "count": 1, | |
| 445 | + | "items": [{ | |
| 446 | + | "@id": format!("{index}#page/{lower}/{upper}"), | |
| 447 | + | "count": versions.len(), | |
| 448 | + | "lower": lower, | |
| 449 | + | "upper": upper, | |
| 450 | + | "items": versions.iter().map(|v| leaf(base, id, v)).collect::<Vec<_>>(), | |
| 451 | + | }], | |
| 452 | + | }) | |
| 453 | + | } | |
| 454 | + | ||
| 455 | + | /// One package as search answers it: its listed versions, the newest as | |
| 456 | + | /// its version. `None` when none is listed. | |
| 457 | + | pub fn search_result(base: &str, id: &str, versions: &[Listed<'_>]) -> Option<Value> { | |
| 458 | + | let shown: Vec<&Listed<'_>> = versions.iter().filter(|v| v.listed).collect(); | |
| 459 | + | let newest = shown.iter().max_by(|a, b| compare(a.version, b.version))?; | |
| 460 | + | let m = newest.metadata; | |
| 461 | + | let at = addresses(base, id, newest.version); | |
| 462 | + | let authors: Vec<&str> = m["authors"].as_str().map(|a| a.split(',').map(str::trim).filter(|a| !a.is_empty()).collect()).unwrap_or_default(); | |
| 463 | + | Some(json!({ | |
| 464 | + | "@id": at.registration, | |
| 465 | + | "@type": "Package", | |
| 466 | + | "registration": at.registration, | |
| 467 | + | "id": m["id"].as_str().unwrap_or(id), | |
| 468 | + | "version": newest.version, | |
| 469 | + | "description": m["description"].as_str().unwrap_or(""), | |
| 470 | + | "summary": m["summary"].as_str().unwrap_or(""), | |
| 471 | + | "title": m["title"].as_str().unwrap_or(""), | |
| 472 | + | "projectUrl": m["project_url"].as_str().unwrap_or(""), | |
| 473 | + | "licenseUrl": m["license_url"].as_str().unwrap_or(""), | |
| 474 | + | "iconUrl": m["icon_url"].as_str().unwrap_or(""), | |
| 475 | + | "authors": authors, | |
| 476 | + | "tags": m["tags"].as_array().cloned().unwrap_or_default(), | |
| 477 | + | "totalDownloads": shown.iter().map(|v| v.downloads).sum::<u64>(), | |
| 478 | + | "verified": false, | |
| 479 | + | "packageTypes": [{ "name": "Dependency" }], | |
| 480 | + | "versions": shown.iter().map(|v| json!({ | |
| 481 | + | "version": v.version, | |
| 482 | + | "downloads": v.downloads, | |
| 483 | + | "@id": addresses(base, id, v.version).leaf, | |
| 484 | + | })).collect::<Vec<_>>(), | |
| 485 | + | })) | |
| 486 | + | } | |
| 487 | + | ||
| 488 | + | #[cfg(test)] | |
| 489 | + | mod tests { | |
| 490 | + | use super::*; | |
| 491 | + | ||
| 492 | + | #[test] | |
| 493 | + | fn ids_follow_nugets_rules() { | |
| 494 | + | for good in ["Acme.Web", "Newtonsoft.Json", "a", "my-lib_2", &"a".repeat(100)] { | |
| 495 | + | assert!(valid_id(good), "{good}"); | |
| 496 | + | } | |
| 497 | + | for bad in ["", ".a", "a.", "a..b", "a b", "a/b", "a.-b", &"a".repeat(101)] { | |
| 498 | + | assert!(!valid_id(bad), "{bad}"); | |
| 499 | + | } | |
| 500 | + | } | |
| 501 | + | ||
| 502 | + | #[test] | |
| 503 | + | fn versions_normalize_and_order_as_nuget_does() { | |
| 504 | + | assert_eq!(normalize("1.0").as_deref(), Some("1.0.0")); | |
| 505 | + | assert_eq!(normalize("1.0.0.0").as_deref(), Some("1.0.0")); | |
| 506 | + | assert_eq!(normalize("1.0.0.4").as_deref(), Some("1.0.0.4")); | |
| 507 | + | assert_eq!(normalize("01.02.3").as_deref(), Some("1.2.3")); | |
| 508 | + | assert_eq!(normalize("1.0.0-Beta.1+sha.abc").as_deref(), Some("1.0.0-Beta.1")); | |
| 509 | + | for bad in ["", "a.b", "1.0.0.0.0", "1..0", "1.0-", "1.0-a..b", "1.0+"] { | |
| 510 | + | assert_eq!(normalize(bad), None, "{bad}"); | |
| 511 | + | } | |
| 512 | + | assert!(is_prerelease("1.0.0-rc.1") && !is_prerelease("1.0.0")); | |
| 513 | + | let mut versions = vec!["1.0.0", "1.0.0-beta.2", "1.0.0-beta.10", "1.0.0-alpha", "0.9.0", "1.0.0.1", "1.0.0-BETA"]; | |
| 514 | + | versions.sort_by(|a, b| compare(a, b)); | |
| 515 | + | assert_eq!(versions, ["0.9.0", "1.0.0-alpha", "1.0.0-BETA", "1.0.0-beta.2", "1.0.0-beta.10", "1.0.0", "1.0.0.1"]); | |
| 516 | + | } | |
| 517 | + | ||
| 518 | + | #[test] | |
| 519 | + | fn every_endpoint_is_routed() { | |
| 520 | + | let at = |route: NugetRoute| Some(("acme".to_owned(), route)); | |
| 521 | + | assert_eq!(route("/-/nuget/Acme/v3/index.json"), at(NugetRoute::Index)); | |
| 522 | + | assert_eq!(route("/-/nuget/acme/v3/query"), at(NugetRoute::Search)); | |
| 523 | + | assert_eq!(route("/-/nuget/acme/v3/flatcontainer/acme.web/index.json"), at(NugetRoute::Versions { id: "acme.web".into() })); | |
| 524 | + | assert_eq!( | |
| 525 | + | route("/-/nuget/acme/v3/flatcontainer/acme.web/1.0.0/acme.web.1.0.0.nupkg"), | |
| 526 | + | at(NugetRoute::Content { id: "acme.web".into(), version: "1.0.0".into(), file: Content::Nupkg }) | |
| 527 | + | ); | |
| 528 | + | assert_eq!( | |
| 529 | + | route("/-/nuget/acme/v3/flatcontainer/acme.web/1.0.0/acme.web.nuspec"), | |
| 530 | + | at(NugetRoute::Content { id: "acme.web".into(), version: "1.0.0".into(), file: Content::Nuspec }) | |
| 531 | + | ); | |
| 532 | + | assert_eq!(route("/-/nuget/acme/v3/flatcontainer/acme.web/1.0.0/other.1.0.0.nupkg"), None); | |
| 533 | + | assert_eq!(route("/-/nuget/acme/v3/registration/acme.web/index.json"), at(NugetRoute::Registration { id: "acme.web".into() })); | |
| 534 | + | assert_eq!(route("/-/nuget/acme/v3/registration/acme.web/1.0.0.json"), at(NugetRoute::Leaf { id: "acme.web".into(), version: "1.0.0".into() })); | |
| 535 | + | assert_eq!(route("/-/nuget/acme/api/v2/package"), at(NugetRoute::Push)); | |
| 536 | + | assert_eq!(route("/-/nuget/acme/api/v2/package/"), at(NugetRoute::Push)); | |
| 537 | + | assert_eq!(route("/-/nuget/acme/api/v2/package/Acme.Web/1.0.0"), at(NugetRoute::Listing { id: "Acme.Web".into(), version: "1.0.0".into() })); | |
| 538 | + | assert_eq!(route("/-/nuget/acme/v3/flatcontainer/a..b/index.json"), None); | |
| 539 | + | assert_eq!(route("/-/nuget/acme"), None); | |
| 540 | + | assert_eq!(route("/-/nuget/acme/v2"), None); | |
| 541 | + | } | |
| 542 | + | ||
| 543 | + | #[test] | |
| 544 | + | fn the_push_body_is_multipart_with_the_package() { | |
| 545 | + | let body = b"--abc123\r\nContent-Type: application/octet-stream\r\nContent-Disposition: form-data; name=package; filename=package.nupkg\r\n\r\nPK\x03\x04data\r\n--abc\r\nmore\r\n--abc123--\r\n"; | |
| 546 | + | assert_eq!(pushed_file(Some("multipart/form-data; boundary=\"abc123\""), body).unwrap(), b"PK\x03\x04data\r\n--abc\r\nmore"); | |
| 547 | + | assert_eq!(pushed_file(Some("application/octet-stream"), b"PK raw").unwrap(), b"PK raw"); | |
| 548 | + | assert_eq!(pushed_file(None, b"PK raw").unwrap(), b"PK raw"); | |
| 549 | + | assert!(pushed_file(Some("multipart/form-data"), body).is_err(), "no boundary"); | |
| 550 | + | assert!(pushed_file(Some("multipart/form-data; boundary=zzz"), body).is_err()); | |
| 551 | + | } | |
| 552 | + | ||
| 553 | + | const NUSPEC: &str = r#"<?xml version="1.0" encoding="utf-8"?> | |
| 554 | + | <package xmlns="http://schemas.microsoft.com/packaging/2013/05/nuspec.xsd"> | |
| 555 | + | <metadata> | |
| 556 | + | <id>Acme.Web</id> | |
| 557 | + | <version>1.2.0</version> | |
| 558 | + | <authors>Ada, Bo</authors> | |
| 559 | + | <description>The web client.</description> | |
| 560 | + | <license type="expression">MIT</license> | |
| 561 | + | <readme>docs\README.md</readme> | |
| 562 | + | <repository type="git" url="https://g1t.sh/acme/web.git" /> | |
| 563 | + | <tags>http client</tags> | |
| 564 | + | <dependencies> | |
| 565 | + | <group targetFramework="net8.0"> | |
| 566 | + | <dependency id="Newtonsoft.Json" version="13.0.1" exclude="Build,Analyzers" /> | |
| 567 | + | <dependency id="Acme.Core" version="[1.0.0, 2.0.0)" /> | |
| 568 | + | </group> | |
| 569 | + | <group targetFramework=".NETStandard2.0" /> | |
| 570 | + | </dependencies> | |
| 571 | + | </metadata> | |
| 572 | + | </package>"#; | |
| 573 | + | ||
| 574 | + | #[test] | |
| 575 | + | fn a_package_is_read_from_its_nuspec() { | |
| 576 | + | let nupkg = crate::composer::zip(&[ | |
| 577 | + | ("Acme.Web.nuspec".to_owned(), NUSPEC.as_bytes().to_vec()), | |
| 578 | + | ("docs/README.md".to_owned(), b"# Acme.Web\n".to_vec()), | |
| 579 | + | ("lib/net8.0/Acme.Web.dll".to_owned(), b"MZ".to_vec()), | |
| 580 | + | ]); | |
| 581 | + | let package = read_package(&nupkg).unwrap(); | |
| 582 | + | let spec = &package.nuspec; | |
| 583 | + | assert_eq!((spec.id.as_str(), spec.version.as_str()), ("Acme.Web", "1.2.0")); | |
| 584 | + | assert_eq!(spec.license_expression.as_deref(), Some("MIT")); | |
| 585 | + | assert_eq!(spec.repository_url.as_deref(), Some("https://g1t.sh/acme/web.git")); | |
| 586 | + | assert_eq!(spec.tags, ["http", "client"]); | |
| 587 | + | assert_eq!(package.readme.as_deref(), Some("# Acme.Web\n")); | |
| 588 | + | assert_eq!(spec.groups.len(), 2); | |
| 589 | + | assert_eq!(spec.groups[0].target_framework.as_deref(), Some("net8.0")); | |
| 590 | + | assert_eq!(spec.groups[0].dependencies, [("Newtonsoft.Json".into(), "[13.0.1, )".into()), ("Acme.Core".into(), "[1.0.0, 2.0.0)".into())]); | |
| 591 | + | assert!(spec.groups[1].dependencies.is_empty()); | |
| 592 | + | assert!(read_package(b"not a zip").is_err()); | |
| 593 | + | let empty = crate::composer::zip(&[("lib/a.dll".to_owned(), b"MZ".to_vec())]); | |
| 594 | + | assert!(read_package(&empty).is_err(), "no .nuspec"); | |
| 595 | + | assert_eq!(range(None), "(, )"); | |
| 596 | + | } | |
| 597 | + | ||
| 598 | + | #[test] | |
| 599 | + | fn the_documents_are_nugets_shape() { | |
| 600 | + | let index = service_index("https://g1t.sh/-/nuget/acme"); | |
| 601 | + | let kinds: Vec<&str> = index["resources"].as_array().unwrap().iter().map(|r| r["@type"].as_str().unwrap()).collect(); | |
| 602 | + | for kind in ["PackageBaseAddress/3.0.0", "RegistrationsBaseUrl", "SearchQueryService", "PackagePublish/2.0.0"] { | |
| 603 | + | assert!(kinds.contains(&kind), "{kind}"); | |
| 604 | + | } | |
| 605 | + | let spec = read_nuspec(NUSPEC).unwrap(); | |
| 606 | + | let (one, two) = (stored(&spec, "1.0.0"), stored(&spec, "1.2.0")); | |
| 607 | + | let versions = [ | |
| 608 | + | Listed { version: "1.0.0", metadata: &one, published: "2026-10-01T00:00:00.000Z", listed: false, downloads: 3 }, | |
| 609 | + | Listed { version: "1.2.0", metadata: &two, published: "2026-10-06T00:00:00.000Z", listed: true, downloads: 4 }, | |
| 610 | + | ]; | |
| 611 | + | let base = "https://g1t.sh/-/nuget/acme"; | |
| 612 | + | let reg = registration(base, "Acme.Web", &versions); | |
| 613 | + | let page = ®["items"][0]; | |
| 614 | + | assert_eq!(page["lower"], "1.0.0"); | |
| 615 | + | assert_eq!(page["upper"], "1.2.0"); | |
| 616 | + | let entry = &page["items"][1]["catalogEntry"]; | |
| 617 | + | assert_eq!(entry["id"], "Acme.Web"); | |
| 618 | + | assert_eq!(entry["listed"], true); | |
| 619 | + | assert_eq!(entry["packageContent"], "https://g1t.sh/-/nuget/acme/v3/flatcontainer/acme.web/1.2.0/acme.web.1.2.0.nupkg"); | |
| 620 | + | assert_eq!(entry["dependencyGroups"][0]["targetFramework"], "net8.0"); | |
| 621 | + | assert_eq!(entry["dependencyGroups"][0]["dependencies"][1]["range"], "[1.0.0, 2.0.0)"); | |
| 622 | + | assert_eq!(page["items"][0]["catalogEntry"]["listed"], false); | |
| 623 | + | let found = search_result(base, "Acme.Web", &versions).unwrap(); | |
| 624 | + | assert_eq!(found["version"], "1.2.0"); | |
| 625 | + | assert_eq!(found["versions"].as_array().unwrap().len(), 1, "unlisted versions are not searched"); | |
| 626 | + | assert_eq!(found["totalDownloads"], 4); | |
| 627 | + | assert_eq!(found["authors"], json!(["Ada", "Bo"])); | |
| 628 | + | assert!(search_result(base, "Acme.Web", &versions[..1]).is_none()); | |
| 629 | + | } | |
| 630 | + | } |
| 1 | + | //! The NuGet feed: `g1t.sh/-/nuget/<workspace>/v3/index.json`, a v3 feed | |
| 2 | + | //! for each workspace. `dotnet nuget push` sends a g1t token as its API key | |
| 3 | + | //! (`X-NuGet-ApiKey`); restores send Basic credentials (any username, a g1t | |
| 4 | + | //! token as the password) from `nuget.config`, after the feed answers a | |
| 5 | + | //! private request with a `401`. | |
| 6 | + | //! | |
| 7 | + | //! A `.nupkg` is stored once, by its SHA-256, with its `.nuspec` beside it; | |
| 8 | + | //! the flat container, registration and search documents are made from the | |
| 9 | + | //! versions on each read. `dotnet nuget delete` unlists a version, as | |
| 10 | + | //! nuget.org does: it is still downloaded by those who name it. | |
| 11 | + | ||
| 12 | + | use g1t_contracts::User; | |
| 13 | + | use g1t_contracts::audit::AuditActor; | |
| 14 | + | use g1t_contracts::events::PackageEvent; | |
| 15 | + | use g1t_contracts::new_id; | |
| 16 | + | use g1t_kit::now_ms; | |
| 17 | + | use serde_json::{Value, json}; | |
| 18 | + | use worker::{Context, Headers, Method, Request, Response, ResponseBody, Result, Url}; | |
| 19 | + | ||
| 20 | + | use crate::access::{self, Action}; | |
| 21 | + | use crate::db::{NewFile, NewVersion, PackageRow, VersionRow}; | |
| 22 | + | use crate::digest::Digest; | |
| 23 | + | use crate::npm; | |
| 24 | + | use crate::nuget::{self, Content, Listed, NugetRoute}; | |
| 25 | + | use crate::oci::{Credentials, origin, published_by}; | |
| 26 | + | use crate::store::BlobStore; | |
| 27 | + | use crate::{Caller, Packages, TargetOf, token}; | |
| 28 | + | ||
| 29 | + | const NUGET: &str = "nuget"; | |
| 30 | + | /// The most versions a package's documents list. | |
| 31 | + | const MAX_VERSIONS: u32 = 5000; | |
| 32 | + | /// The longest README kept for a package's page. | |
| 33 | + | const MAX_README_BYTES: usize = 1024 * 1024; | |
| 34 | + | /// The most packages one search answers with. | |
| 35 | + | const MAX_SEARCH: u32 = 100; | |
| 36 | + | const DOCS: &str = "https://docs.g1t.sh/guides/nuget/"; | |
| 37 | + | const TOKENS: &str = "https://g1t.sh/settings/tokens"; | |
| 38 | + | ||
| 39 | + | /// A plain-text answer, which `dotnet` prints after the status. | |
| 40 | + | fn error(status: u16, message: impl Into<String>) -> Result<Response> { | |
| 41 | + | let mut response = Response::ok(message.into())?.with_status(status); | |
| 42 | + | response.headers_mut().set("content-type", "text/plain; charset=utf-8")?; | |
| 43 | + | if status == 401 { | |
| 44 | + | response.headers_mut().set("www-authenticate", "Basic realm=\"g1t\"")?; | |
| 45 | + | } | |
| 46 | + | Ok(response) | |
| 47 | + | } | |
| 48 | + | ||
| 49 | + | fn sign_in() -> String { | |
| 50 | + | format!("Sign in to use this feed: give the source a username and a g1t access token from {TOKENS} as its password. See {DOCS}") | |
| 51 | + | } | |
| 52 | + | ||
| 53 | + | fn json_response(value: &Value, head: bool) -> Result<Response> { | |
| 54 | + | let mut response = if head { Response::empty()? } else { Response::from_json(value)? }; | |
| 55 | + | response.headers_mut().set("content-type", "application/json")?; | |
| 56 | + | response.headers_mut().set("cache-control", "no-cache")?; | |
| 57 | + | Ok(response) | |
| 58 | + | } | |
| 59 | + | ||
| 60 | + | impl Packages { | |
| 61 | + | /// Answers a NuGet request. | |
| 62 | + | pub async fn nuget(&self, request: Request, ctx: &Context) -> Result<Response> { | |
| 63 | + | let url = request.url()?; | |
| 64 | + | let Some((workspace, route)) = nuget::route(url.path()) else { | |
| 65 | + | return error(404, "There is nothing at this address."); | |
| 66 | + | }; | |
| 67 | + | match self.nuget_route(request, &url, &workspace, route, ctx).await { | |
| 68 | + | Ok(response) => Ok(response), | |
| 69 | + | Err(problem) => { | |
| 70 | + | worker::console_error!("packages: nuget {}: {problem}", url.path()); | |
| 71 | + | error(500, "Something went wrong on our side. Try again in a moment.") | |
| 72 | + | } | |
| 73 | + | } | |
| 74 | + | } | |
| 75 | + | ||
| 76 | + | /// Who the request is from: the push's API key, or Basic credentials | |
| 77 | + | /// (or a `Bearer` token) from the source's settings. | |
| 78 | + | async fn nuget_credentials(&self, request: &Request) -> Result<Credentials> { | |
| 79 | + | if let Some(key) = request.headers().get("x-nuget-apikey")?.map(|k| k.trim().to_owned()).filter(|k| !k.is_empty()) { | |
| 80 | + | return Ok(match self.viewer_for("token", &key).await? { | |
| 81 | + | Some(user) => Credentials::Viewer(Some(user)), | |
| 82 | + | None => Credentials::Bad, | |
| 83 | + | }); | |
| 84 | + | } | |
| 85 | + | let Some(header) = request.headers().get("authorization")? else { | |
| 86 | + | return Ok(Credentials::None); | |
| 87 | + | }; | |
| 88 | + | let viewer = if let Some((username, secret)) = token::basic(&header) { | |
| 89 | + | self.viewer_for(&username, &secret).await? | |
| 90 | + | } else if let Some(bearer) = token::bearer(&header) { | |
| 91 | + | self.viewer_for("token", bearer).await? | |
| 92 | + | } else { | |
| 93 | + | None | |
| 94 | + | }; | |
| 95 | + | Ok(match viewer { | |
| 96 | + | Some(user) => Credentials::Viewer(Some(user)), | |
| 97 | + | None => Credentials::Bad, | |
| 98 | + | }) | |
| 99 | + | } | |
| 100 | + | ||
| 101 | + | async fn nuget_route(&self, mut request: Request, url: &Url, workspace: &str, route: NugetRoute, ctx: &Context) -> Result<Response> { | |
| 102 | + | let method = request.method(); | |
| 103 | + | let credentials = self.nuget_credentials(&request).await?; | |
| 104 | + | let read = matches!(method, Method::Get | Method::Head); | |
| 105 | + | if read && let Some(refused) = self.limited(&request, &credentials, "a g1t token in the source's credentials").await? { | |
| 106 | + | return Ok(refused); | |
| 107 | + | } | |
| 108 | + | let viewer = match credentials { | |
| 109 | + | Credentials::Viewer(viewer) => viewer, | |
| 110 | + | Credentials::None => None, | |
| 111 | + | Credentials::Token(_) | Credentials::Bad => { | |
| 112 | + | return error(401, format!("The token is not right, or has expired. Make an access token at {TOKENS}.")); | |
| 113 | + | } | |
| 114 | + | }; | |
| 115 | + | let viewer = viewer.as_ref(); | |
| 116 | + | let base = format!("{}/-/nuget/{workspace}", origin(url)); | |
| 117 | + | let head = method == Method::Head; | |
| 118 | + | match route { | |
| 119 | + | NugetRoute::Index if read => { | |
| 120 | + | if viewer.is_none() && self.db.has_private(workspace, NUGET).await? { | |
| 121 | + | return error(401, sign_in()); | |
| 122 | + | } | |
| 123 | + | json_response(&nuget::service_index(&base), head) | |
| 124 | + | } | |
| 125 | + | NugetRoute::Versions { id } if read => self.nuget_versions(workspace, &id, viewer, head).await, | |
| 126 | + | NugetRoute::Content { id, version, file } if read => self.nuget_content(workspace, &id, &version, file, viewer, head, ctx).await, | |
| 127 | + | NugetRoute::Registration { id } if read => self.nuget_registration(&base, workspace, &id, None, viewer, head).await, | |
| 128 | + | NugetRoute::Leaf { id, version } if read => self.nuget_registration(&base, workspace, &id, Some(&version), viewer, head).await, | |
| 129 | + | NugetRoute::Search if read => self.nuget_search(url, &base, workspace, viewer).await, | |
| 130 | + | NugetRoute::Push if method == Method::Put => self.nuget_push(&mut request, workspace, viewer).await, | |
| 131 | + | NugetRoute::Listing { id, version } if method == Method::Delete => self.nuget_listing(workspace, &id, &version, false, viewer).await, | |
| 132 | + | NugetRoute::Listing { id, version } if method == Method::Post => self.nuget_listing(workspace, &id, &version, true, viewer).await, | |
| 133 | + | _ => error(405, "Not a method this address takes."), | |
| 134 | + | } | |
| 135 | + | } | |
| 136 | + | ||
| 137 | + | /// The package, by its id in any case, if its workspace is not deleted. | |
| 138 | + | async fn nuget_package(&self, workspace: &str, id: &str) -> Result<Option<PackageRow>> { | |
| 139 | + | Ok(self.db.package_any_case(workspace, NUGET, id).await?.filter(|p| !p.hidden())) | |
| 140 | + | } | |
| 141 | + | ||
| 142 | + | /// The answer for something not there: a `401` to someone not signed | |
| 143 | + | /// in when the workspace has private packages, so the client sends its | |
| 144 | + | /// credentials and a private package looks like a missing one. | |
| 145 | + | async fn nuget_absent(&self, workspace: &str, viewer: Option<&User>) -> Result<Response> { | |
| 146 | + | if viewer.is_none() && self.db.has_private(workspace, NUGET).await? { | |
| 147 | + | return error(401, sign_in()); | |
| 148 | + | } | |
| 149 | + | error(404, "Not found: no such package or version, or you cannot see it.") | |
| 150 | + | } | |
| 151 | + | ||
| 152 | + | async fn nuget_check(&self, viewer: Option<&User>, package: &PackageRow, action: Action) -> Result<Option<Response>> { | |
| 153 | + | let target = TargetOf::package(package); | |
| 154 | + | let decision = access::decide(viewer, &target.view(), action); | |
| 155 | + | if decision.allowed { | |
| 156 | + | return Ok(None); | |
| 157 | + | } | |
| 158 | + | let readable = action != Action::Pull && access::decide(viewer, &target.view(), Action::Pull).allowed; | |
| 159 | + | if !readable && viewer.is_none() { | |
| 160 | + | return Ok(Some(error(401, sign_in())?)); | |
| 161 | + | } | |
| 162 | + | if !readable { | |
| 163 | + | return Ok(Some(self.nuget_absent(&package.workspace, viewer).await?)); | |
| 164 | + | } | |
| 165 | + | Ok(Some(error(403, decision.reason.unwrap_or_else(|| "Not allowed.".to_owned()))?)) | |
| 166 | + | } | |
| 167 | + | ||
| 168 | + | /// The package and its versions, oldest first, when the viewer may read it. | |
| 169 | + | async fn nuget_readable(&self, workspace: &str, id: &str, viewer: Option<&User>) -> Result<std::result::Result<(PackageRow, Vec<VersionRow>), Response>> { | |
| 170 | + | let Some(package) = self.nuget_package(workspace, id).await? else { | |
| 171 | + | return Ok(Err(self.nuget_absent(workspace, viewer).await?)); | |
| 172 | + | }; | |
| 173 | + | if let Some(refusal) = self.nuget_check(viewer, &package, Action::Pull).await? { | |
| 174 | + | return Ok(Err(refusal)); | |
| 175 | + | } | |
| 176 | + | let mut versions = self.db.versions(&package.id, MAX_VERSIONS).await?; | |
| 177 | + | if versions.is_empty() { | |
| 178 | + | return Ok(Err(self.nuget_absent(workspace, viewer).await?)); | |
| 179 | + | } | |
| 180 | + | versions.sort_by(|a, b| nuget::compare(&a.version, &b.version)); | |
| 181 | + | Ok(Ok((package, versions))) | |
| 182 | + | } | |
| 183 | + | ||
| 184 | + | /// The flat container's version list: every version, unlisted ones too. | |
| 185 | + | async fn nuget_versions(&self, workspace: &str, id: &str, viewer: Option<&User>, head: bool) -> Result<Response> { | |
| 186 | + | let (_, versions) = match self.nuget_readable(workspace, id, viewer).await? { | |
| 187 | + | Ok(found) => found, | |
| 188 | + | Err(refused) => return Ok(refused), | |
| 189 | + | }; | |
| 190 | + | let listed: Vec<String> = versions.iter().map(|v| v.version.to_ascii_lowercase()).collect(); | |
| 191 | + | json_response(&json!({ "versions": listed }), head) | |
| 192 | + | } | |
| 193 | + | ||
| 194 | + | /// A version's `.nupkg` or `.nuspec`. | |
| 195 | + | #[allow(clippy::too_many_arguments)] | |
| 196 | + | async fn nuget_content(&self, workspace: &str, id: &str, version: &str, file: Content, viewer: Option<&User>, head: bool, ctx: &Context) -> Result<Response> { | |
| 197 | + | let (package, versions) = match self.nuget_readable(workspace, id, viewer).await? { | |
| 198 | + | Ok(found) => found, | |
| 199 | + | Err(refused) => return Ok(refused), | |
| 200 | + | }; | |
| 201 | + | let wanted = nuget::normalize(version).unwrap_or_default().to_ascii_lowercase(); | |
| 202 | + | let Some(row) = versions.iter().find(|v| v.version.to_ascii_lowercase() == wanted) else { | |
| 203 | + | return self.nuget_absent(workspace, viewer).await; | |
| 204 | + | }; | |
| 205 | + | let name = match file { | |
| 206 | + | Content::Nupkg => "nupkg", | |
| 207 | + | Content::Nuspec => "nuspec", | |
| 208 | + | }; | |
| 209 | + | let Some(kept) = self.db.file(&row.id, name).await? else { | |
| 210 | + | return self.nuget_absent(workspace, viewer).await; | |
| 211 | + | }; | |
| 212 | + | let Some(digest) = Digest::parse(&kept.digest) else { | |
| 213 | + | return self.nuget_absent(workspace, viewer).await; | |
| 214 | + | }; | |
| 215 | + | let Some(blob) = self.db.package_blob(&package.id, &digest).await? else { | |
| 216 | + | return self.nuget_absent(workspace, viewer).await; | |
| 217 | + | }; | |
| 218 | + | let headers = Headers::new(); | |
| 219 | + | headers.set("content-type", if file == Content::Nupkg { "application/octet-stream" } else { "application/xml" })?; | |
| 220 | + | headers.set("content-length", &blob.size.to_string())?; | |
| 221 | + | headers.set("cache-control", "max-age=31536000")?; | |
| 222 | + | if head { | |
| 223 | + | return Ok(Response::from_body(ResponseBody::Empty)?.with_headers(headers)); | |
| 224 | + | } | |
| 225 | + | let Some(got) = self.store.get(&blob.object_key, None).await? else { | |
| 226 | + | return self.nuget_absent(workspace, viewer).await; | |
| 227 | + | }; | |
| 228 | + | if file == Content::Nupkg { | |
| 229 | + | self.count_download(&package.id, ctx); | |
| 230 | + | } | |
| 231 | + | Ok(Response::from_body(got.body)?.with_headers(headers)) | |
| 232 | + | } | |
| 233 | + | ||
| 234 | + | /// A package's registration index, or one version's leaf. | |
| 235 | + | async fn nuget_registration(&self, base: &str, workspace: &str, id: &str, version: Option<&str>, viewer: Option<&User>, head: bool) -> Result<Response> { | |
| 236 | + | let (package, versions) = match self.nuget_readable(workspace, id, viewer).await? { | |
| 237 | + | Ok(found) => found, | |
| 238 | + | Err(refused) => return Ok(refused), | |
| 239 | + | }; | |
| 240 | + | let metadata: Vec<Value> = versions.iter().map(VersionRow::meta).collect(); | |
| 241 | + | let listed: Vec<Listed<'_>> = versions | |
| 242 | + | .iter() | |
| 243 | + | .zip(&metadata) | |
| 244 | + | .map(|(row, metadata)| Listed { version: &row.version, metadata, published: &row.published_at, listed: !row.is_yanked(), downloads: 0 }) | |
| 245 | + | .collect(); | |
| 246 | + | match version { | |
| 247 | + | None => json_response(&nuget::registration(base, &package.name, &listed), head), | |
| 248 | + | Some(version) => { | |
| 249 | + | let wanted = nuget::normalize(version).unwrap_or_default().to_ascii_lowercase(); | |
| 250 | + | let Some(one) = listed.iter().find(|v| v.version.to_ascii_lowercase() == wanted) else { | |
| 251 | + | return self.nuget_absent(workspace, viewer).await; | |
| 252 | + | }; | |
| 253 | + | json_response(&nuget::leaf(base, &package.name, one), head) | |
| 254 | + | } | |
| 255 | + | } | |
| 256 | + | } | |
| 257 | + | ||
| 258 | + | /// Search: the workspace's packages the viewer may see whose id or | |
| 259 | + | /// description holds the query, with their listed versions. | |
| 260 | + | async fn nuget_search(&self, url: &Url, base: &str, workspace: &str, viewer: Option<&User>) -> Result<Response> { | |
| 261 | + | let query = |key: &str| url.query_pairs().find(|(k, _)| k == key).map(|(_, v)| v.into_owned()); | |
| 262 | + | let q = query("q").unwrap_or_default().trim().to_ascii_lowercase(); | |
| 263 | + | let skip = query("skip").and_then(|v| v.parse::<usize>().ok()).unwrap_or(0); | |
| 264 | + | let take = query("take").and_then(|v| v.parse::<usize>().ok()).unwrap_or(20).min(MAX_SEARCH as usize); | |
| 265 | + | let prerelease = query("prerelease").is_some_and(|v| v.eq_ignore_ascii_case("true")); | |
| 266 | + | if viewer.is_none() && self.db.has_private(workspace, NUGET).await? { | |
| 267 | + | return error(401, sign_in()); | |
| 268 | + | } | |
| 269 | + | let packages = self.db.packages_of(workspace, NUGET, 1000).await?; | |
| 270 | + | let versions = self.db.ecosystem_versions(workspace, NUGET, 20_000).await?; | |
| 271 | + | let mut found = Vec::new(); | |
| 272 | + | for package in &packages { | |
| 273 | + | if !access::decide(viewer, &TargetOf::package(package).view(), Action::Pull).allowed { | |
| 274 | + | continue; | |
| 275 | + | } | |
| 276 | + | let rows: Vec<&VersionRow> = versions | |
| 277 | + | .iter() | |
| 278 | + | .filter(|v| v.package_id == package.id && (prerelease || !nuget::is_prerelease(&v.version))) | |
| 279 | + | .collect(); | |
| 280 | + | let metadata: Vec<Value> = rows.iter().map(|v| v.meta()).collect(); | |
| 281 | + | let matches = q.is_empty() | |
| 282 | + | || package.name.to_ascii_lowercase().contains(&q) | |
| 283 | + | || package.description.as_deref().is_some_and(|d| d.to_ascii_lowercase().contains(&q)); | |
| 284 | + | if !matches { | |
| 285 | + | continue; | |
| 286 | + | } | |
| 287 | + | let mut listed: Vec<Listed<'_>> = rows | |
| 288 | + | .iter() | |
| 289 | + | .zip(&metadata) | |
| 290 | + | .map(|(row, metadata)| Listed { version: &row.version, metadata, published: &row.published_at, listed: !row.is_yanked(), downloads: 0 }) | |
| 291 | + | .collect(); | |
| 292 | + | listed.sort_by(|a, b| nuget::compare(a.version, b.version)); | |
| 293 | + | if let Some(mut result) = nuget::search_result(base, &package.name, &listed) { | |
| 294 | + | result["totalDownloads"] = json!(package.downloads); | |
| 295 | + | found.push(result); | |
| 296 | + | } | |
| 297 | + | } | |
| 298 | + | let total = found.len(); | |
| 299 | + | let data: Vec<Value> = found.into_iter().skip(skip).take(take).collect(); | |
| 300 | + | json_response(&json!({ "totalHits": total, "data": data }), false) | |
| 301 | + | } | |
| 302 | + | ||
| 303 | + | /// The package a first push makes, linked to the repository its | |
| 304 | + | /// `.nuspec` names on g1t, or else one named like its id. | |
| 305 | + | async fn nuget_target(&self, workspace: &str, id: &str, repository: Option<&str>) -> Result<TargetOf> { | |
| 306 | + | let named = repository | |
| 307 | + | .and_then(|url| npm::repository_of(&Value::String(url.to_owned()), &self.host)) | |
| 308 | + | .filter(|(owner, _)| owner == workspace) | |
| 309 | + | .map(|(_, repo)| repo); | |
| 310 | + | let lower = id.to_ascii_lowercase(); | |
| 311 | + | let dashed = lower.replace('.', "-"); | |
| 312 | + | let mut repo = None; | |
| 313 | + | for candidate in named.iter().map(String::as_str).chain([lower.as_str(), dashed.as_str()]) { | |
| 314 | + | if let Some(found) = self.repo_by_name(workspace, candidate).await? { | |
| 315 | + | repo = Some(found); | |
| 316 | + | break; | |
| 317 | + | } | |
| 318 | + | } | |
| 319 | + | Ok(TargetOf { workspace: workspace.to_owned(), repo: repo.map(|r| (r.id, r.name, r.is_private)), public: false }) | |
| 320 | + | } | |
| 321 | + | ||
| 322 | + | /// `dotnet nuget push`: a `PUT` of the `.nupkg`, in a multipart body. | |
| 323 | + | async fn nuget_push(&self, request: &mut Request, workspace: &str, viewer: Option<&User>) -> Result<Response> { | |
| 324 | + | let declared = request.headers().get("content-length")?.and_then(|n| n.parse::<u64>().ok()); | |
| 325 | + | let too_large = || { | |
| 326 | + | let mb = self.max_request / 1_000_000; | |
| 327 | + | error(413, format!("A push may be at most {mb} MB. See {DOCS}#size")) | |
| 328 | + | }; | |
| 329 | + | if declared.is_some_and(|n| n > self.max_request) { | |
| 330 | + | return too_large(); | |
| 331 | + | } | |
| 332 | + | let content_type = request.headers().get("content-type")?; | |
| 333 | + | let body = request.bytes().await?; | |
| 334 | + | if body.len() as u64 > self.max_request { | |
| 335 | + | return too_large(); | |
| 336 | + | } | |
| 337 | + | if viewer.is_none() { | |
| 338 | + | return error(401, format!("Push with a g1t access token as the API key: dotnet nuget push <file> --api-key <token>. Make one at {TOKENS}.")); | |
| 339 | + | } | |
| 340 | + | let nupkg = match nuget::pushed_file(content_type.as_deref(), &body) { | |
| 341 | + | Ok(file) => file, | |
| 342 | + | Err(message) => return error(400, message), | |
| 343 | + | }; | |
| 344 | + | let read = match nuget::read_package(nupkg) { | |
| 345 | + | Ok(read) => read, | |
| 346 | + | Err(message) => return error(400, message), | |
| 347 | + | }; | |
| 348 | + | let spec = &read.nuspec; | |
| 349 | + | if !nuget::valid_id(&spec.id) { | |
| 350 | + | return error(400, format!("{} is not a valid package id: letters, digits and _, in parts joined by ., - or _.", spec.id)); | |
| 351 | + | } | |
| 352 | + | let Some(version) = nuget::normalize(&spec.version) else { | |
| 353 | + | return error(400, format!("{} is not a version NuGet reads.", spec.version)); | |
| 354 | + | }; | |
| 355 | + | ||
| 356 | + | let found = self.db.package_any_case(workspace, NUGET, &spec.id).await?; | |
| 357 | + | if found.as_ref().is_some_and(PackageRow::hidden) || (found.is_none() && self.db.workspace_hidden(workspace).await?) { | |
| 358 | + | return error(403, format!("The workspace {workspace} is deleted; nothing can be pushed to it.")); | |
| 359 | + | } | |
| 360 | + | let target = match &found { | |
| 361 | + | Some(package) => TargetOf::package(package), | |
| 362 | + | None => self.nuget_target(workspace, &spec.id, spec.repository_url.as_deref()).await?, | |
| 363 | + | }; | |
| 364 | + | let decision = access::decide(viewer, &target.view(), Action::Push); | |
| 365 | + | if !decision.allowed { | |
| 366 | + | let readable = found.is_none() || access::decide(viewer, &target.view(), Action::Pull).allowed; | |
| 367 | + | if !readable { | |
| 368 | + | return error(404, "Not found: no such package, or you cannot see it."); | |
| 369 | + | } | |
| 370 | + | return error(403, decision.reason.unwrap_or_else(|| "Not allowed.".to_owned())); | |
| 371 | + | } | |
| 372 | + | let caller = Caller { actor: viewer.map(AuditActor::of) }; | |
| 373 | + | let package = match found { | |
| 374 | + | Some(package) => package, | |
| 375 | + | None => { | |
| 376 | + | self.db | |
| 377 | + | .create_package( | |
| 378 | + | &new_id("pkg", now_ms()), | |
| 379 | + | workspace, | |
| 380 | + | NUGET, | |
| 381 | + | &spec.id, | |
| 382 | + | target.repo.as_ref().map(|(id, repo, private)| (id.as_str(), repo.as_str(), *private)), | |
| 383 | + | caller.actor.as_ref().map_or("", |actor| actor.actor_id.as_str()), | |
| 384 | + | now_ms(), | |
| 385 | + | ) | |
| 386 | + | .await? | |
| 387 | + | } | |
| 388 | + | }; | |
| 389 | + | let existing = self.db.versions(&package.id, MAX_VERSIONS).await?; | |
| 390 | + | if let Some(taken) = existing.iter().find(|v| v.version.eq_ignore_ascii_case(&version)) { | |
| 391 | + | return error(409, format!("{} {} is already pushed, and a version is pushed once. Bump the version.", package.name, taken.version)); | |
| 392 | + | } | |
| 393 | + | ||
| 394 | + | let nupkg = nupkg.to_vec(); | |
| 395 | + | let digest = Digest::of(&nupkg); | |
| 396 | + | let size = nupkg.len() as u64; | |
| 397 | + | let nuspec_digest = Digest::of(&read.nuspec_bytes); | |
| 398 | + | let nuspec_size = read.nuspec_bytes.len() as u64; | |
| 399 | + | let files = [(digest.to_string(), size), (nuspec_digest.to_string(), nuspec_size)]; | |
| 400 | + | if let Some(refusal) = self.storage_refusal(&package, &files).await? { | |
| 401 | + | return error(403, refusal); | |
| 402 | + | } | |
| 403 | + | let now = now_ms(); | |
| 404 | + | for (digest, bytes, media_type) in [(&digest, nupkg, "application/octet-stream"), (&nuspec_digest, read.nuspec_bytes.clone(), "application/xml")] { | |
| 405 | + | let stored = match self.db.blob(digest).await? { | |
| 406 | + | Some(blob) => self.store.head(&blob.object_key).await?.is_some(), | |
| 407 | + | None => false, | |
| 408 | + | }; | |
| 409 | + | let length = bytes.len() as u64; | |
| 410 | + | if !stored { | |
| 411 | + | self.store.put(&digest.object_key(), bytes).await?; | |
| 412 | + | } | |
| 413 | + | self.db.keep_blob(&package.id, digest, length, Some(media_type), &digest.object_key(), now).await?; | |
| 414 | + | } | |
| 415 | + | self.db | |
| 416 | + | .publish( | |
| 417 | + | NewVersion { | |
| 418 | + | id: new_id("ver", now), | |
| 419 | + | package_id: package.id.clone(), | |
| 420 | + | version: version.clone(), | |
| 421 | + | digest: digest.to_string(), | |
| 422 | + | size: size + nuspec_size, | |
| 423 | + | metadata: nuget::stored(spec, &version).to_string(), | |
| 424 | + | subject: None, | |
| 425 | + | published_by: published_by(&caller), | |
| 426 | + | files: vec![ | |
| 427 | + | NewFile { name: "nupkg".to_owned(), digest: digest.to_string(), size, media_type: Some("application/octet-stream".to_owned()) }, | |
| 428 | + | NewFile { | |
| 429 | + | name: "nuspec".to_owned(), | |
| 430 | + | digest: nuspec_digest.to_string(), | |
| 431 | + | size: nuspec_size, | |
| 432 | + | media_type: Some("application/xml".to_owned()), | |
| 433 | + | }, | |
| 434 | + | ], | |
| 435 | + | }, | |
| 436 | + | None, | |
| 437 | + | now, | |
| 438 | + | ) | |
| 439 | + | .await?; | |
| 440 | + | // The README and description the page shows: the highest stable | |
| 441 | + | // version's, so a pre-release does not replace them. | |
| 442 | + | let highest = !nuget::is_prerelease(&version) | |
| 443 | + | && existing.iter().filter(|v| !nuget::is_prerelease(&v.version)).all(|v| nuget::compare(&v.version, &version).is_lt()); | |
| 444 | + | let first = nuget::is_prerelease(&version) && existing.is_empty(); | |
| 445 | + | if highest || first { | |
| 446 | + | let readme = read.readme.as_deref().map(str::trim).filter(|r| !r.is_empty() && r.len() <= MAX_README_BYTES); | |
| 447 | + | let readme_digest = match readme { | |
| 448 | + | Some(readme) => { | |
| 449 | + | let bytes = readme.as_bytes().to_vec(); | |
| 450 | + | let digest = Digest::of(&bytes); | |
| 451 | + | if self.db.blob(&digest).await?.is_none() { | |
| 452 | + | self.store.put(&digest.object_key(), bytes.clone()).await?; | |
| 453 | + | } | |
| 454 | + | self.db.keep_blob(&package.id, &digest, bytes.len() as u64, Some("text/markdown"), &digest.object_key(), now).await?; | |
| 455 | + | Some(digest.to_string()) | |
| 456 | + | } | |
| 457 | + | None => None, | |
| 458 | + | }; | |
| 459 | + | self.db.set_readme(&package.id, readme_digest.as_deref(), spec.description.as_deref(), now).await?; | |
| 460 | + | } | |
| 461 | + | self.db.measure(&package.workspace).await?; | |
| 462 | + | let event = PackageEvent { | |
| 463 | + | version: Some(version.clone()), | |
| 464 | + | digest: Some(digest.to_string()), | |
| 465 | + | size: Some(size), | |
| 466 | + | ..self.event_of(&package) | |
| 467 | + | }; | |
| 468 | + | self.announce("package.published", &package, event, &caller).await; | |
| 469 | + | self.audit(&caller, "package.publish", &package, Some(&format!("{workspace}/{}@{version}", package.name)), None).await; | |
| 470 | + | error(201, format!("{} {version} was pushed.", package.name)) | |
| 471 | + | } | |
| 472 | + | ||
| 473 | + | /// `dotnet nuget delete` unlists a version; a `POST` lists it again. | |
| 474 | + | async fn nuget_listing(&self, workspace: &str, id: &str, version: &str, listed: bool, viewer: Option<&User>) -> Result<Response> { | |
| 475 | + | let Some(package) = self.nuget_package(workspace, id).await? else { | |
| 476 | + | return self.nuget_absent(workspace, viewer).await; | |
| 477 | + | }; | |
| 478 | + | if let Some(refusal) = self.nuget_check(viewer, &package, Action::Push).await? { | |
| 479 | + | return Ok(refusal); | |
| 480 | + | } | |
| 481 | + | let wanted = nuget::normalize(version).unwrap_or_default(); | |
| 482 | + | let versions = self.db.versions(&package.id, MAX_VERSIONS).await?; | |
| 483 | + | let Some(row) = versions.iter().find(|v| v.version.eq_ignore_ascii_case(&wanted)) else { | |
| 484 | + | return error(404, format!("{} {version} is not there.", package.name)); | |
| 485 | + | }; | |
| 486 | + | if row.is_yanked() == listed { | |
| 487 | + | self.db.set_yanked(&row.id, !listed).await?; | |
| 488 | + | self.db.touch_package(&package.id, now_ms()).await?; | |
| 489 | + | let caller = Caller { actor: viewer.map(AuditActor::of) }; | |
| 490 | + | let action = if listed { "package.relist" } else { "package.unlist" }; | |
| 491 | + | self.audit(&caller, action, &package, Some(&format!("{workspace}/{}@{}", package.name, row.version)), None).await; | |
| 492 | + | } | |
| 493 | + | Ok(Response::empty()?.with_status(if listed { 200 } else { 204 })) | |
| 494 | + | } | |
| 495 | + | } |
| 975 | 975 | ..self.event_of(package) | |
| 976 | 976 | }; | |
| 977 | 977 | self.announce("package.version_deleted", package, event, caller).await; | |
| 978 | − | // npm and Cargo name a version by its number; an image by its digest. | |
| 978 | + | // An image (and a Composer version, by its commit) is named by its | |
| 979 | + | // digest; every other package by its version. | |
| 979 | 980 | let path = if package.ecosystem == "npm" { | |
| 980 | 981 | format!("@{}/{}@{}", package.workspace, package.name, version.version) | |
| 981 | − | } else if package.ecosystem == "cargo" { | |
| 982 | + | } else if !matches!(package.ecosystem.as_str(), "container" | "composer") { | |
| 982 | 983 | format!("{}/{}@{}", package.workspace, package.name, version.version) | |
| 983 | 984 | } else { | |
| 984 | 985 | format!("{}/{}@{}", package.workspace, package.name, version.digest) |
| 1 | + | //! What the RubyGems registry needs that does not touch the network: gem | |
| 2 | + | //! names and versions, the registry's paths, the `Gem::Specification` read | |
| 3 | + | //! from a `.gem` (a tar holding `metadata.gz`), and the compact index | |
| 4 | + | //! Bundler reads: `versions`, `info/<gem>` and `names`. | |
| 5 | + | //! | |
| 6 | + | //! A version is keyed by its number and platform as the compact index | |
| 7 | + | //! writes it (`1.0.0`, `1.0.0-x86_64-linux`), and keeps what its index | |
| 8 | + | //! line needs as its metadata, made once when it is pushed. | |
| 9 | + | ||
| 10 | + | use serde_json::{Value, json}; | |
| 11 | + | ||
| 12 | + | use crate::archive; | |
| 13 | + | use crate::yaml; | |
| 14 | + | ||
| 15 | + | /// The longest gem name taken. | |
| 16 | + | pub const MAX_NAME: usize = 128; | |
| 17 | + | /// The largest `metadata.gz`, unpacked, read from a gem. | |
| 18 | + | const MAX_METADATA_BYTES: usize = 4 * 1024 * 1024; | |
| 19 | + | ||
| 20 | + | /// A gem's name: letters, digits, `.`, `-` and `_`, with a letter in it. | |
| 21 | + | pub fn valid_name(name: &str) -> bool { | |
| 22 | + | !name.is_empty() | |
| 23 | + | && name.len() <= MAX_NAME | |
| 24 | + | && name.bytes().all(|b| b.is_ascii_alphanumeric() || matches!(b, b'.' | b'-' | b'_')) | |
| 25 | + | && name.bytes().any(|b| b.is_ascii_alphabetic()) | |
| 26 | + | && name.as_bytes()[0].is_ascii_alphanumeric() | |
| 27 | + | } | |
| 28 | + | ||
| 29 | + | /// A version as `Gem::Version` takes it: numbers and words joined by dots, | |
| 30 | + | /// starting with a number (`1.0.0`, `2.0.0.rc1`, `1.0.0-beta.1`). | |
| 31 | + | pub fn valid_version(version: &str) -> bool { | |
| 32 | + | let (core, suffix) = version.split_once('-').map_or((version, None), |(c, s)| (c, Some(s))); | |
| 33 | + | let mut parts = core.split('.'); | |
| 34 | + | let first_ok = parts.next().is_some_and(|p| !p.is_empty() && p.bytes().all(|b| b.is_ascii_digit())); | |
| 35 | + | first_ok | |
| 36 | + | && version.len() <= 128 | |
| 37 | + | && parts.all(|p| !p.is_empty() && p.bytes().all(|b| b.is_ascii_alphanumeric())) | |
| 38 | + | && suffix.is_none_or(|s| s.split('.').all(|p| !p.is_empty() && p.bytes().all(|b| b.is_ascii_alphanumeric() || b == b'-'))) | |
| 39 | + | } | |
| 40 | + | ||
| 41 | + | /// A version with a letter in it is a pre-release, as RubyGems decides. | |
| 42 | + | pub fn is_prerelease(version: &str) -> bool { | |
| 43 | + | version.bytes().any(|b| b.is_ascii_alphabetic()) | |
| 44 | + | } | |
| 45 | + | ||
| 46 | + | /// The version as the index keys it: `1.0.0`, or `1.0.0-java` for a gem | |
| 47 | + | /// built for a platform. | |
| 48 | + | pub fn key(version: &str, platform: &str) -> String { | |
| 49 | + | if platform.is_empty() || platform == "ruby" { version.to_owned() } else { format!("{version}-{platform}") } | |
| 50 | + | } | |
| 51 | + | ||
| 52 | + | /// One of the registry's endpoints, under `/-/rubygems/<workspace>/`. | |
| 53 | + | #[derive(Clone, Debug, PartialEq, Eq)] | |
| 54 | + | pub enum GemRoute { | |
| 55 | + | /// `versions`: every gem and its versions, for Bundler. | |
| 56 | + | Versions, | |
| 57 | + | /// `info/<gem>`: a gem's versions, dependencies and checksums. | |
| 58 | + | Info { name: String }, | |
| 59 | + | /// `names`: every gem's name. | |
| 60 | + | Names, | |
| 61 | + | /// `gems/<name>-<version>[-<platform>].gem`; the name and version are | |
| 62 | + | /// told apart by the handler, as names may hold `-`. | |
| 63 | + | Gem { stem: String }, | |
| 64 | + | /// `api/v1/gems`: `gem push`. | |
| 65 | + | Push, | |
| 66 | + | /// `api/v1/gems/yank`: `gem yank`. | |
| 67 | + | Yank, | |
| 68 | + | } | |
| 69 | + | ||
| 70 | + | pub fn route(path: &str) -> Option<(String, GemRoute)> { | |
| 71 | + | let rest = path.strip_prefix("/-/rubygems/")?; | |
| 72 | + | let (workspace, rest) = rest.split_once('/')?; | |
| 73 | + | let workspace = workspace.to_ascii_lowercase(); | |
| 74 | + | if workspace.is_empty() { | |
| 75 | + | return None; | |
| 76 | + | } | |
| 77 | + | let route = match rest.trim_end_matches('/') { | |
| 78 | + | "versions" => GemRoute::Versions, | |
| 79 | + | "names" => GemRoute::Names, | |
| 80 | + | "api/v1/gems" => GemRoute::Push, | |
| 81 | + | "api/v1/gems/yank" => GemRoute::Yank, | |
| 82 | + | other => { | |
| 83 | + | if let Some(name) = other.strip_prefix("info/") { | |
| 84 | + | if !valid_name(name) { | |
| 85 | + | return None; | |
| 86 | + | } | |
| 87 | + | GemRoute::Info { name: name.to_owned() } | |
| 88 | + | } else { | |
| 89 | + | let file = other.strip_prefix("gems/")?; | |
| 90 | + | let stem = file.strip_suffix(".gem")?; | |
| 91 | + | if stem.contains('/') || candidates(stem).is_empty() { | |
| 92 | + | return None; | |
| 93 | + | } | |
| 94 | + | GemRoute::Gem { stem: stem.to_owned() } | |
| 95 | + | } | |
| 96 | + | } | |
| 97 | + | }; | |
| 98 | + | Some((workspace, route)) | |
| 99 | + | } | |
| 100 | + | ||
| 101 | + | /// The ways a file's stem splits into a name and a version key: at each | |
| 102 | + | /// `-` followed by a digit, longest name last (`a-b-1.0` is `a-b` `1.0`). | |
| 103 | + | pub fn candidates(stem: &str) -> Vec<(String, String)> { | |
| 104 | + | stem.char_indices() | |
| 105 | + | .filter(|&(i, c)| c == '-' && stem[i + 1..].starts_with(|n: char| n.is_ascii_digit())) | |
| 106 | + | .map(|(i, _)| (stem[..i].to_owned(), stem[i + 1..].to_owned())) | |
| 107 | + | .filter(|(name, _)| valid_name(name)) | |
| 108 | + | .collect() | |
| 109 | + | } | |
| 110 | + | ||
| 111 | + | /// What is read from a gem's specification. | |
| 112 | + | #[derive(Clone, Debug, Default, PartialEq, Eq)] | |
| 113 | + | pub struct Gemspec { | |
| 114 | + | pub name: String, | |
| 115 | + | pub version: String, | |
| 116 | + | pub platform: String, | |
| 117 | + | pub summary: Option<String>, | |
| 118 | + | pub description: Option<String>, | |
| 119 | + | pub homepage: Option<String>, | |
| 120 | + | pub source_code_uri: Option<String>, | |
| 121 | + | pub licenses: Vec<String>, | |
| 122 | + | pub authors: Vec<String>, | |
| 123 | + | /// Runtime dependencies: name and requirement (`>= 2.0&< 4`). | |
| 124 | + | pub dependencies: Vec<(String, String)>, | |
| 125 | + | pub ruby: Option<String>, | |
| 126 | + | pub rubygems: Option<String>, | |
| 127 | + | } | |
| 128 | + | ||
| 129 | + | /// A `Gem::Requirement` as the compact index writes it: `>= 2.0&< 4`. | |
| 130 | + | pub fn requirement(value: &Value) -> String { | |
| 131 | + | value["requirements"] | |
| 132 | + | .as_array() | |
| 133 | + | .map(|list| { | |
| 134 | + | list.iter() | |
| 135 | + | .filter_map(|pair| { | |
| 136 | + | let op = pair.get(0)?.as_str()?; | |
| 137 | + | let version = pair.get(1).map(|v| v.get("version").unwrap_or(v)).and_then(Value::as_str)?; | |
| 138 | + | Some(format!("{op} {version}")) | |
| 139 | + | }) | |
| 140 | + | .collect::<Vec<_>>() | |
| 141 | + | .join("&") | |
| 142 | + | }) | |
| 143 | + | .unwrap_or_default() | |
| 144 | + | } | |
| 145 | + | ||
| 146 | + | /// A requirement that says nothing (`>= 0`), as left out of the index. | |
| 147 | + | fn anything(requirement: &str) -> bool { | |
| 148 | + | requirement.is_empty() || requirement == ">= 0" | |
| 149 | + | } | |
| 150 | + | ||
| 151 | + | fn text(value: &Value) -> Option<String> { | |
| 152 | + | value.as_str().map(str::trim).filter(|s| !s.is_empty()).map(str::to_owned) | |
| 153 | + | } | |
| 154 | + | ||
| 155 | + | fn texts(value: &Value) -> Vec<String> { | |
| 156 | + | match value { | |
| 157 | + | Value::Array(list) => list.iter().filter_map(text).collect(), | |
| 158 | + | Value::String(_) => text(value).into_iter().collect(), | |
| 159 | + | _ => Vec::new(), | |
| 160 | + | } | |
| 161 | + | } | |
| 162 | + | ||
| 163 | + | /// Reads a specification RubyGems wrote as YAML. | |
| 164 | + | pub fn read_spec(text_yaml: &str) -> Result<Gemspec, String> { | |
| 165 | + | let spec = yaml::parse(text_yaml).map_err(|problem| format!("The gem's metadata is not YAML: {problem}"))?; | |
| 166 | + | if !spec.is_object() { | |
| 167 | + | return Err("The gem's metadata is not a specification.".to_owned()); | |
| 168 | + | } | |
| 169 | + | let version = &spec["version"]; | |
| 170 | + | let version = text(version.get("version").unwrap_or(version)).ok_or("The gem's metadata names no version.")?; | |
| 171 | + | let platform = match &spec["platform"] { | |
| 172 | + | Value::Object(map) => ["cpu", "os", "version"].iter().filter_map(|k| map.get(*k).and_then(text)).collect::<Vec<_>>().join("-"), | |
| 173 | + | other => text(other).unwrap_or_else(|| "ruby".to_owned()), | |
| 174 | + | }; | |
| 175 | + | let dependencies = spec["dependencies"] | |
| 176 | + | .as_array() | |
| 177 | + | .map(|deps| { | |
| 178 | + | deps.iter() | |
| 179 | + | .filter(|d| d["type"].as_str().is_none_or(|t| t.trim_start_matches(':') == "runtime")) | |
| 180 | + | .filter_map(|d| { | |
| 181 | + | let name = text(&d["name"])?; | |
| 182 | + | let req = d.get("requirement").filter(|r| r.is_object()).or_else(|| d.get("version_requirements")).map(requirement).unwrap_or_default(); | |
| 183 | + | Some((name, if req.is_empty() { ">= 0".to_owned() } else { req })) | |
| 184 | + | }) | |
| 185 | + | .collect() | |
| 186 | + | }) | |
| 187 | + | .unwrap_or_default(); | |
| 188 | + | let required = |key: &str| Some(requirement(&spec[key])).filter(|r| !anything(r)); | |
| 189 | + | Ok(Gemspec { | |
| 190 | + | name: text(&spec["name"]).ok_or("The gem's metadata names no gem.")?, | |
| 191 | + | version, | |
| 192 | + | platform: if platform.is_empty() { "ruby".to_owned() } else { platform }, | |
| 193 | + | summary: text(&spec["summary"]), | |
| 194 | + | description: text(&spec["description"]), | |
| 195 | + | homepage: text(&spec["homepage"]), | |
| 196 | + | source_code_uri: text(&spec["metadata"]["source_code_uri"]), | |
| 197 | + | licenses: texts(&spec["licenses"]), | |
| 198 | + | authors: texts(&spec["authors"]), | |
| 199 | + | dependencies, | |
| 200 | + | ruby: required("required_ruby_version"), | |
| 201 | + | rubygems: required("required_rubygems_version"), | |
| 202 | + | }) | |
| 203 | + | } | |
| 204 | + | ||
| 205 | + | /// Reads a `.gem`: the specification in its `metadata.gz`. | |
| 206 | + | pub fn read_gem(gem: &[u8]) -> Result<Gemspec, String> { | |
| 207 | + | let files = archive::tar_files(gem).map_err(|_| "The file is not a .gem: it is not a tar archive.".to_owned())?; | |
| 208 | + | let (_, metadata) = files.iter().find(|(name, _)| name == "metadata.gz").ok_or("The gem has no metadata.gz.")?; | |
| 209 | + | let yaml = archive::gunzip(metadata, MAX_METADATA_BYTES)?; | |
| 210 | + | read_spec(&String::from_utf8(yaml).map_err(|_| "The gem's metadata is not UTF-8.".to_owned())?) | |
| 211 | + | } | |
| 212 | + | ||
| 213 | + | /// What a version keeps for its index line and the gem's page. | |
| 214 | + | pub fn stored(spec: &Gemspec) -> Value { | |
| 215 | + | json!({ | |
| 216 | + | "name": spec.name, | |
| 217 | + | "number": spec.version, | |
| 218 | + | "platform": spec.platform, | |
| 219 | + | "summary": spec.summary, | |
| 220 | + | "description": spec.description, | |
| 221 | + | "homepage": spec.homepage, | |
| 222 | + | "source_code_uri": spec.source_code_uri, | |
| 223 | + | "licenses": spec.licenses, | |
| 224 | + | "authors": spec.authors, | |
| 225 | + | "dependencies": spec.dependencies.iter().map(|(name, req)| json!({ "name": name, "requirement": req })).collect::<Vec<_>>(), | |
| 226 | + | "ruby": spec.ruby, | |
| 227 | + | "rubygems": spec.rubygems, | |
| 228 | + | }) | |
| 229 | + | } | |
| 230 | + | ||
| 231 | + | /// One line of a gem's info file: `1.0.0 rack:>= 2.0&< 4|checksum:<sha256>,ruby:>= 3.0`. | |
| 232 | + | pub fn info_line(key: &str, stored: &Value, checksum: &str) -> String { | |
| 233 | + | let deps: Vec<String> = stored["dependencies"] | |
| 234 | + | .as_array() | |
| 235 | + | .map(|deps| { | |
| 236 | + | deps.iter() | |
| 237 | + | .filter_map(|d| Some(format!("{}:{}", d["name"].as_str()?, d["requirement"].as_str().unwrap_or(">= 0")))) | |
| 238 | + | .collect() | |
| 239 | + | }) | |
| 240 | + | .unwrap_or_default(); | |
| 241 | + | let mut requirements = vec![format!("checksum:{checksum}")]; | |
| 242 | + | if let Some(ruby) = stored["ruby"].as_str() { | |
| 243 | + | requirements.push(format!("ruby:{ruby}")); | |
| 244 | + | } | |
| 245 | + | if let Some(rubygems) = stored["rubygems"].as_str() { | |
| 246 | + | requirements.push(format!("rubygems:{rubygems}")); | |
| 247 | + | } | |
| 248 | + | format!("{key} {}|{}", deps.join(","), requirements.join(",")) | |
| 249 | + | } | |
| 250 | + | ||
| 251 | + | /// A gem's info file, from its versions' lines, oldest first. | |
| 252 | + | pub fn info(lines: &[String]) -> String { | |
| 253 | + | let mut out = String::from("---\n"); | |
| 254 | + | for line in lines { | |
| 255 | + | out.push_str(line); | |
| 256 | + | out.push('\n'); | |
| 257 | + | } | |
| 258 | + | out | |
| 259 | + | } | |
| 260 | + | ||
| 261 | + | /// The `versions` file: each gem's versions and its info file's MD5. | |
| 262 | + | pub fn versions_file(created_at: &str, gems: &[(String, Vec<String>, String)]) -> String { | |
| 263 | + | let mut out = format!("created_at: {created_at}\n---\n"); | |
| 264 | + | for (name, versions, md5) in gems { | |
| 265 | + | out.push_str(&format!("{name} {} {md5}\n", versions.join(","))); | |
| 266 | + | } | |
| 267 | + | out | |
| 268 | + | } | |
| 269 | + | ||
| 270 | + | pub fn names_file(names: &[String]) -> String { | |
| 271 | + | let mut out = String::from("---\n"); | |
| 272 | + | for name in names { | |
| 273 | + | out.push_str(name); | |
| 274 | + | out.push('\n'); | |
| 275 | + | } | |
| 276 | + | out | |
| 277 | + | } | |
| 278 | + | ||
| 279 | + | /// A value from a form body or query string (`gem_name=hello&version=1.0`). | |
| 280 | + | pub fn form_value(form: &str, key: &str) -> Option<String> { | |
| 281 | + | let url = worker::Url::parse(&format!("http://form.invalid/?{form}")).ok()?; | |
| 282 | + | url.query_pairs().find(|(k, _)| k == key).map(|(_, v)| v.into_owned()).filter(|v| !v.is_empty()) | |
| 283 | + | } | |
| 284 | + | ||
| 285 | + | #[cfg(test)] | |
| 286 | + | mod tests { | |
| 287 | + | use super::*; | |
| 288 | + | ||
| 289 | + | #[test] | |
| 290 | + | fn names_and_versions_follow_rubygems_rules() { | |
| 291 | + | for good in ["rails", "hello-world", "net_http2", "a1", "Hello.rb"] { | |
| 292 | + | assert!(valid_name(good), "{good}"); | |
| 293 | + | } | |
| 294 | + | for bad in ["", "123", "-a", ".a", "a b", "a/b", &"a".repeat(129)] { | |
| 295 | + | assert!(!valid_name(bad), "{bad}"); | |
| 296 | + | } | |
| 297 | + | for good in ["1.0.0", "0.1", "2.0.0.rc1", "1.0.0-beta.1", "3"] { | |
| 298 | + | assert!(valid_version(good), "{good}"); | |
| 299 | + | } | |
| 300 | + | for bad in ["", "a.1", "1..0", "1.0 0", "1.0-"] { | |
| 301 | + | assert!(!valid_version(bad), "{bad}"); | |
| 302 | + | } | |
| 303 | + | assert!(is_prerelease("2.0.0.rc1") && !is_prerelease("2.0.0")); | |
| 304 | + | assert_eq!(key("1.0.0", "ruby"), "1.0.0"); | |
| 305 | + | assert_eq!(key("1.0.0", "x86_64-linux"), "1.0.0-x86_64-linux"); | |
| 306 | + | } | |
| 307 | + | ||
| 308 | + | #[test] | |
| 309 | + | fn every_endpoint_is_routed() { | |
| 310 | + | let at = |route: GemRoute| Some(("acme".to_owned(), route)); | |
| 311 | + | assert_eq!(route("/-/rubygems/Acme/versions"), at(GemRoute::Versions)); | |
| 312 | + | assert_eq!(route("/-/rubygems/acme/names"), at(GemRoute::Names)); | |
| 313 | + | assert_eq!(route("/-/rubygems/acme/info/hello-world"), at(GemRoute::Info { name: "hello-world".into() })); | |
| 314 | + | assert_eq!(route("/-/rubygems/acme/gems/hello-world-0.1.0.gem"), at(GemRoute::Gem { stem: "hello-world-0.1.0".into() })); | |
| 315 | + | assert_eq!(route("/-/rubygems/acme/api/v1/gems"), at(GemRoute::Push)); | |
| 316 | + | assert_eq!(route("/-/rubygems/acme/api/v1/gems/yank"), at(GemRoute::Yank)); | |
| 317 | + | assert_eq!(route("/-/rubygems/acme/gems/hello.gem"), None, "no version"); | |
| 318 | + | assert_eq!(route("/-/rubygems/acme/info/a b"), None); | |
| 319 | + | assert_eq!(route("/-/rubygems/acme/other"), None); | |
| 320 | + | assert_eq!(route("/-/rubygems/acme"), None); | |
| 321 | + | assert_eq!( | |
| 322 | + | candidates("hello-world-0.1.0-x86_64-linux"), | |
| 323 | + | [("hello-world".to_owned(), "0.1.0-x86_64-linux".to_owned())], | |
| 324 | + | "x86_64 starts with a letter" | |
| 325 | + | ); | |
| 326 | + | assert_eq!(candidates("a-2-1.0"), [("a".to_owned(), "2-1.0".to_owned()), ("a-2".to_owned(), "1.0".to_owned())]); | |
| 327 | + | } | |
| 328 | + | ||
| 329 | + | const SPEC: &str = r#"--- !ruby/object:Gem::Specification | |
| 330 | + | name: hello-world | |
| 331 | + | version: !ruby/object:Gem::Version | |
| 332 | + | version: 0.2.0 | |
| 333 | + | platform: ruby | |
| 334 | + | authors: | |
| 335 | + | - Ada | |
| 336 | + | dependencies: | |
| 337 | + | - !ruby/object:Gem::Dependency | |
| 338 | + | name: rack | |
| 339 | + | requirement: !ruby/object:Gem::Requirement | |
| 340 | + | requirements: | |
| 341 | + | - - "<" | |
| 342 | + | - !ruby/object:Gem::Version | |
| 343 | + | version: '4' | |
| 344 | + | - - ">=" | |
| 345 | + | - !ruby/object:Gem::Version | |
| 346 | + | version: '2.0' | |
| 347 | + | type: :runtime | |
| 348 | + | prerelease: false | |
| 349 | + | version_requirements: !ruby/object:Gem::Requirement | |
| 350 | + | requirements: | |
| 351 | + | - - "<" | |
| 352 | + | - !ruby/object:Gem::Version | |
| 353 | + | version: '4' | |
| 354 | + | - !ruby/object:Gem::Dependency | |
| 355 | + | name: json | |
| 356 | + | requirement: !ruby/object:Gem::Requirement | |
| 357 | + | requirements: | |
| 358 | + | - - ">=" | |
| 359 | + | - !ruby/object:Gem::Version | |
| 360 | + | version: '0' | |
| 361 | + | type: :runtime | |
| 362 | + | - !ruby/object:Gem::Dependency | |
| 363 | + | name: rspec | |
| 364 | + | requirement: !ruby/object:Gem::Requirement | |
| 365 | + | requirements: | |
| 366 | + | - - "~>" | |
| 367 | + | - !ruby/object:Gem::Version | |
| 368 | + | version: '3.0' | |
| 369 | + | type: :development | |
| 370 | + | description: Says hello. | |
| 371 | + | homepage: https://g1t.sh/acme/hello-world | |
| 372 | + | licenses: | |
| 373 | + | - MIT | |
| 374 | + | metadata: | |
| 375 | + | source_code_uri: https://g1t.sh/acme/hello-world | |
| 376 | + | required_ruby_version: !ruby/object:Gem::Requirement | |
| 377 | + | requirements: | |
| 378 | + | - - ">=" | |
| 379 | + | - !ruby/object:Gem::Version | |
| 380 | + | version: 3.0.0 | |
| 381 | + | required_rubygems_version: !ruby/object:Gem::Requirement | |
| 382 | + | requirements: | |
| 383 | + | - - ">=" | |
| 384 | + | - !ruby/object:Gem::Version | |
| 385 | + | version: '0' | |
| 386 | + | summary: Says hello | |
| 387 | + | "#; | |
| 388 | + | ||
| 389 | + | #[test] | |
| 390 | + | fn a_gem_is_read_from_its_metadata() { | |
| 391 | + | let mut gz = vec![0x1f, 0x8b, 8, 0, 0, 0, 0, 0, 0, 3]; | |
| 392 | + | gz.extend_from_slice(&miniz_oxide::deflate::compress_to_vec(SPEC.as_bytes(), 6)); | |
| 393 | + | gz.extend_from_slice(&crate::composer::crc32(SPEC.as_bytes()).to_le_bytes()); | |
| 394 | + | gz.extend_from_slice(&(SPEC.len() as u32).to_le_bytes()); | |
| 395 | + | let mut gem = Vec::new(); | |
| 396 | + | for (name, data) in [("metadata.gz", gz.as_slice()), ("data.tar.gz", b"x".as_slice())] { | |
| 397 | + | let mut header = [0u8; 512]; | |
| 398 | + | header[..name.len()].copy_from_slice(name.as_bytes()); | |
| 399 | + | header[124..136].copy_from_slice(format!("{:011o}\0", data.len()).as_bytes()); | |
| 400 | + | header[156] = b'0'; | |
| 401 | + | gem.extend_from_slice(&header); | |
| 402 | + | gem.extend_from_slice(data); | |
| 403 | + | gem.resize(gem.len().div_ceil(512) * 512, 0); | |
| 404 | + | } | |
| 405 | + | gem.extend_from_slice(&[0; 1024]); | |
| 406 | + | let spec = read_gem(&gem).unwrap(); | |
| 407 | + | assert_eq!((spec.name.as_str(), spec.version.as_str(), spec.platform.as_str()), ("hello-world", "0.2.0", "ruby")); | |
| 408 | + | assert_eq!(spec.dependencies, [("rack".to_owned(), "< 4&>= 2.0".to_owned()), ("json".to_owned(), ">= 0".to_owned())], "runtime only"); | |
| 409 | + | assert_eq!(spec.ruby.as_deref(), Some(">= 3.0.0")); | |
| 410 | + | assert_eq!(spec.rubygems, None, ">= 0 says nothing"); | |
| 411 | + | assert_eq!(spec.source_code_uri.as_deref(), Some("https://g1t.sh/acme/hello-world")); | |
| 412 | + | assert_eq!(spec.licenses, ["MIT"]); | |
| 413 | + | assert!(read_gem(b"not a gem").is_err()); | |
| 414 | + | ||
| 415 | + | let line = info_line(&key(&spec.version, &spec.platform), &stored(&spec), "abc123"); | |
| 416 | + | assert_eq!(line, "0.2.0 rack:< 4&>= 2.0,json:>= 0|checksum:abc123,ruby:>= 3.0.0"); | |
| 417 | + | let bare = info_line("1.0.0", &json!({ "dependencies": [] }), "ff"); | |
| 418 | + | assert_eq!(bare, "1.0.0 |checksum:ff"); | |
| 419 | + | } | |
| 420 | + | ||
| 421 | + | #[test] | |
| 422 | + | fn a_platform_mapping_reads_as_its_name() { | |
| 423 | + | let spec = read_spec("name: native\nversion: !ruby/object:Gem::Version\n version: 1.0.0\nplatform: !ruby/object:Gem::Platform\n cpu: x86_64\n os: linux\n version:\n").unwrap(); | |
| 424 | + | assert_eq!(spec.platform, "x86_64-linux"); | |
| 425 | + | assert!(read_spec("name: x\n").is_err(), "no version"); | |
| 426 | + | } | |
| 427 | + | ||
| 428 | + | #[test] | |
| 429 | + | fn the_compact_index_is_bundlers_shape() { | |
| 430 | + | let info = info(&["0.1.0 |checksum:aa".to_owned(), "0.2.0 rack:>= 2|checksum:bb".to_owned()]); | |
| 431 | + | assert_eq!(info, "---\n0.1.0 |checksum:aa\n0.2.0 rack:>= 2|checksum:bb\n"); | |
| 432 | + | let versions = versions_file("2026-10-06T00:00:00Z", &[("hello".into(), vec!["0.1.0".into(), "0.2.0".into()], "d41d8".into())]); | |
| 433 | + | assert_eq!(versions, "created_at: 2026-10-06T00:00:00Z\n---\nhello 0.1.0,0.2.0 d41d8\n"); | |
| 434 | + | assert_eq!(names_file(&["a".into(), "b".into()]), "---\na\nb\n"); | |
| 435 | + | assert_eq!(form_value("gem_name=hello-world&version=0.1.0&platform=", "gem_name").as_deref(), Some("hello-world")); | |
| 436 | + | assert_eq!(form_value("gem_name=a%2Bb", "gem_name").as_deref(), Some("a+b")); | |
| 437 | + | assert_eq!(form_value("version=1", "platform"), None); | |
| 438 | + | } | |
| 439 | + | } |
| 1 | + | //! The RubyGems registry: `g1t.sh/-/rubygems/<workspace>/`, one for each | |
| 2 | + | //! workspace. `gem push --host` sends a g1t token as the whole | |
| 3 | + | //! `Authorization` header (`GEM_HOST_API_KEY`, or `~/.gem/credentials`); | |
| 4 | + | //! Bundler sends Basic credentials (any username, a g1t token as the | |
| 5 | + | //! password) from `bundle config`. | |
| 6 | + | //! | |
| 7 | + | //! Bundler installs from the compact index (`versions`, `info/<gem>`, | |
| 8 | + | //! `names`), made from the versions on each read, with each file's MD5 as | |
| 9 | + | //! its `ETag` as Bundler checks it. A `.gem` is stored once, by its | |
| 10 | + | //! SHA-256, which is also its index `checksum`. `gem yank` takes a version | |
| 11 | + | //! out of the index; its file stays for lockfiles that name it. | |
| 12 | + | ||
| 13 | + | use g1t_contracts::User; | |
| 14 | + | use g1t_contracts::audit::AuditActor; | |
| 15 | + | use g1t_contracts::events::PackageEvent; | |
| 16 | + | use g1t_contracts::new_id; | |
| 17 | + | use g1t_kit::now_ms; | |
| 18 | + | use serde_json::Value; | |
| 19 | + | use worker::{Context, Headers, Method, Request, Response, ResponseBody, Result, Url}; | |
| 20 | + | ||
| 21 | + | use crate::access::{self, Action}; | |
| 22 | + | use crate::db::{NewFile, NewVersion, PackageRow, VersionRow}; | |
| 23 | + | use crate::digest::Digest; | |
| 24 | + | use crate::oci::{Credentials, published_by}; | |
| 25 | + | use crate::rubygems::{self, GemRoute}; | |
| 26 | + | use crate::store::BlobStore; | |
| 27 | + | use crate::{Caller, Packages, TargetOf, cargo, npm, token}; | |
| 28 | + | ||
| 29 | + | const RUBYGEMS: &str = "rubygems"; | |
| 30 | + | /// The most versions a workspace's index lists. | |
| 31 | + | const MAX_VERSIONS: u32 = 20_000; | |
| 32 | + | /// The most gems a workspace's index lists. | |
| 33 | + | const MAX_GEMS: u32 = 2000; | |
| 34 | + | const DOCS: &str = "https://docs.g1t.sh/guides/rubygems/"; | |
| 35 | + | const TOKENS: &str = "https://g1t.sh/settings/tokens"; | |
| 36 | + | ||
| 37 | + | /// A plain-text answer, which `gem` and Bundler print. | |
| 38 | + | fn error(status: u16, message: impl Into<String>) -> Result<Response> { | |
| 39 | + | let mut response = Response::ok(message.into())?.with_status(status); | |
| 40 | + | response.headers_mut().set("content-type", "text/plain; charset=utf-8")?; | |
| 41 | + | if status == 401 { | |
| 42 | + | response.headers_mut().set("www-authenticate", "Basic realm=\"g1t\"")?; | |
| 43 | + | } | |
| 44 | + | Ok(response) | |
| 45 | + | } | |
| 46 | + | ||
| 47 | + | fn sign_in(workspace: &str) -> String { | |
| 48 | + | format!( | |
| 49 | + | "Sign in to use this registry: bundle config set --global https://g1t.sh/-/rubygems/{workspace}/ <you>:<token>, with a g1t access token from {TOKENS}. See {DOCS}" | |
| 50 | + | ) | |
| 51 | + | } | |
| 52 | + | ||
| 53 | + | /// An index file, with the quoted MD5 of its body as its `ETag`: Bundler | |
| 54 | + | /// checks the file it keeps against it, and asks with `If-None-Match`. | |
| 55 | + | fn index_file(request: &Request, body: String, head: bool) -> Result<Response> { | |
| 56 | + | let etag = format!("\"{:x}\"", md5::compute(body.as_bytes())); | |
| 57 | + | let headers = Headers::new(); | |
| 58 | + | headers.set("content-type", "text/plain; charset=utf-8")?; | |
| 59 | + | headers.set("cache-control", "no-cache")?; | |
| 60 | + | headers.set("etag", &etag)?; | |
| 61 | + | if request.headers().get("if-none-match")?.is_some_and(|sent| sent.trim_start_matches("W/") == etag) { | |
| 62 | + | return Ok(Response::empty()?.with_status(304).with_headers(headers)); | |
| 63 | + | } | |
| 64 | + | headers.set("content-length", &body.len().to_string())?; | |
| 65 | + | let body = if head { ResponseBody::Empty } else { ResponseBody::Body(body.into_bytes()) }; | |
| 66 | + | Ok(Response::from_body(body)?.with_headers(headers)) | |
| 67 | + | } | |
| 68 | + | ||
| 69 | + | /// A version's line in the index. | |
| 70 | + | fn line(row: &VersionRow) -> String { | |
| 71 | + | let checksum = Digest::parse(&row.digest).map(|d| d.hex().to_owned()).unwrap_or_default(); | |
| 72 | + | rubygems::info_line(&row.version, &row.meta(), &checksum) | |
| 73 | + | } | |
| 74 | + | ||
| 75 | + | impl Packages { | |
| 76 | + | /// Answers a RubyGems request. | |
| 77 | + | pub async fn rubygems(&self, request: Request, ctx: &Context) -> Result<Response> { | |
| 78 | + | let url = request.url()?; | |
| 79 | + | let Some((workspace, route)) = rubygems::route(url.path()) else { | |
| 80 | + | return error(404, "There is nothing at this address."); | |
| 81 | + | }; | |
| 82 | + | match self.rubygems_route(request, &url, &workspace, route, ctx).await { | |
| 83 | + | Ok(response) => Ok(response), | |
| 84 | + | Err(problem) => { | |
| 85 | + | worker::console_error!("packages: rubygems {}: {problem}", url.path()); | |
| 86 | + | error(500, "Something went wrong on our side. Try again in a moment.") | |
| 87 | + | } | |
| 88 | + | } | |
| 89 | + | } | |
| 90 | + | ||
| 91 | + | /// Who the request is from: `gem`'s bare key (or a `Bearer` one), or | |
| 92 | + | /// Bundler's Basic credentials. | |
| 93 | + | async fn gem_credentials(&self, request: &Request) -> Result<Credentials> { | |
| 94 | + | let Some(header) = request.headers().get("authorization")? else { | |
| 95 | + | return Ok(Credentials::None); | |
| 96 | + | }; | |
| 97 | + | let viewer = if let Some((username, secret)) = token::basic(&header) { | |
| 98 | + | self.viewer_for(&username, &secret).await? | |
| 99 | + | } else if let Some(key) = cargo::token(&header) { | |
| 100 | + | self.viewer_for("token", key).await? | |
| 101 | + | } else { | |
| 102 | + | None | |
| 103 | + | }; | |
| 104 | + | Ok(match viewer { | |
| 105 | + | Some(user) => Credentials::Viewer(Some(user)), | |
| 106 | + | None => Credentials::Bad, | |
| 107 | + | }) | |
| 108 | + | } | |
| 109 | + | ||
| 110 | + | async fn rubygems_route(&self, mut request: Request, url: &Url, workspace: &str, route: GemRoute, ctx: &Context) -> Result<Response> { | |
| 111 | + | let method = request.method(); | |
| 112 | + | let credentials = self.gem_credentials(&request).await?; | |
| 113 | + | let read = matches!(method, Method::Get | Method::Head); | |
| 114 | + | if read && let Some(refused) = self.limited(&request, &credentials, "Basic credentials in bundle config").await? { | |
| 115 | + | return Ok(refused); | |
| 116 | + | } | |
| 117 | + | let viewer = match credentials { | |
| 118 | + | Credentials::Viewer(viewer) => viewer, | |
| 119 | + | Credentials::None => None, | |
| 120 | + | Credentials::Token(_) | Credentials::Bad => { | |
| 121 | + | return error(401, format!("The token is not right, or has expired. Make an access token at {TOKENS}.")); | |
| 122 | + | } | |
| 123 | + | }; | |
| 124 | + | let viewer = viewer.as_ref(); | |
| 125 | + | let head = method == Method::Head; | |
| 126 | + | match route { | |
| 127 | + | GemRoute::Versions if read => self.gem_versions(&request, workspace, viewer, head).await, | |
| 128 | + | GemRoute::Names if read => self.gem_names(&request, workspace, viewer, head).await, | |
| 129 | + | GemRoute::Info { name } if read => self.gem_info(&request, workspace, &name, viewer, head).await, | |
| 130 | + | GemRoute::Gem { stem } if read => self.gem_download(workspace, &stem, viewer, head, ctx).await, | |
| 131 | + | GemRoute::Push if method == Method::Post => self.gem_push(&mut request, workspace, viewer).await, | |
| 132 | + | GemRoute::Yank if method == Method::Delete => self.gem_yank(&mut request, url, workspace, viewer).await, | |
| 133 | + | _ => error(405, "Not a method this address takes."), | |
| 134 | + | } | |
| 135 | + | } | |
| 136 | + | ||
| 137 | + | /// The answer for something not there: a `401` to someone not signed | |
| 138 | + | /// in when the workspace has private gems, so Bundler asks for | |
| 139 | + | /// credentials and a private gem looks like a missing one. | |
| 140 | + | async fn gem_absent(&self, workspace: &str, viewer: Option<&User>) -> Result<Response> { | |
| 141 | + | if viewer.is_none() && self.db.has_private(workspace, RUBYGEMS).await? { | |
| 142 | + | return error(401, sign_in(workspace)); | |
| 143 | + | } | |
| 144 | + | error(404, "Not found: no such gem or version, or you cannot see it.") | |
| 145 | + | } | |
| 146 | + | ||
| 147 | + | async fn gem_check(&self, viewer: Option<&User>, package: &PackageRow, action: Action) -> Result<Option<Response>> { | |
| 148 | + | let target = TargetOf::package(package); | |
| 149 | + | let decision = access::decide(viewer, &target.view(), action); | |
| 150 | + | if decision.allowed { | |
| 151 | + | return Ok(None); | |
| 152 | + | } | |
| 153 | + | let readable = action != Action::Pull && access::decide(viewer, &target.view(), Action::Pull).allowed; | |
| 154 | + | if !readable && viewer.is_none() { | |
| 155 | + | return Ok(Some(error(401, sign_in(&package.workspace))?)); | |
| 156 | + | } | |
| 157 | + | if !readable { | |
| 158 | + | return Ok(Some(self.gem_absent(&package.workspace, viewer).await?)); | |
| 159 | + | } | |
| 160 | + | Ok(Some(error(403, decision.reason.unwrap_or_else(|| "Not allowed.".to_owned()))?)) | |
| 161 | + | } | |
| 162 | + | ||
| 163 | + | /// The workspace's gems the viewer may see, with their versions in the | |
| 164 | + | /// index (not yanked), oldest first. | |
| 165 | + | async fn gem_index(&self, workspace: &str, viewer: Option<&User>) -> Result<std::result::Result<Vec<(PackageRow, Vec<VersionRow>)>, Response>> { | |
| 166 | + | if viewer.is_none() && self.db.has_private(workspace, RUBYGEMS).await? { | |
| 167 | + | return Ok(Err(error(401, sign_in(workspace))?)); | |
| 168 | + | } | |
| 169 | + | let packages = self.db.packages_of(workspace, RUBYGEMS, MAX_GEMS).await?; | |
| 170 | + | let versions = self.db.ecosystem_versions(workspace, RUBYGEMS, MAX_VERSIONS).await?; | |
| 171 | + | Ok(Ok(packages | |
| 172 | + | .into_iter() | |
| 173 | + | .filter(|p| access::decide(viewer, &TargetOf::package(p).view(), Action::Pull).allowed) | |
| 174 | + | .map(|p| { | |
| 175 | + | let rows: Vec<VersionRow> = versions.iter().filter(|v| v.package_id == p.id && !v.is_yanked()).cloned().collect(); | |
| 176 | + | (p, rows) | |
| 177 | + | }) | |
| 178 | + | .filter(|(_, rows)| !rows.is_empty()) | |
| 179 | + | .collect())) | |
| 180 | + | } | |
| 181 | + | ||
| 182 | + | async fn gem_versions(&self, request: &Request, workspace: &str, viewer: Option<&User>, head: bool) -> Result<Response> { | |
| 183 | + | let gems = match self.gem_index(workspace, viewer).await? { | |
| 184 | + | Ok(gems) => gems, | |
| 185 | + | Err(refused) => return Ok(refused), | |
| 186 | + | }; | |
| 187 | + | let created = gems | |
| 188 | + | .iter() | |
| 189 | + | .flat_map(|(_, rows)| rows.iter().map(|r| r.published_at.as_str())) | |
| 190 | + | .min() | |
| 191 | + | .map(|at| format!("{}Z", &at[..at.len().min(19)])) | |
| 192 | + | .unwrap_or_else(|| "2026-01-01T00:00:00Z".to_owned()); | |
| 193 | + | let lines: Vec<(String, Vec<String>, String)> = gems | |
| 194 | + | .iter() | |
| 195 | + | .map(|(package, rows)| { | |
| 196 | + | let info = rubygems::info(&rows.iter().map(line).collect::<Vec<_>>()); | |
| 197 | + | (package.name.clone(), rows.iter().map(|r| r.version.clone()).collect(), format!("{:x}", md5::compute(info.as_bytes()))) | |
| 198 | + | }) | |
| 199 | + | .collect(); | |
| 200 | + | index_file(request, rubygems::versions_file(&created, &lines), head) | |
| 201 | + | } | |
| 202 | + | ||
| 203 | + | async fn gem_names(&self, request: &Request, workspace: &str, viewer: Option<&User>, head: bool) -> Result<Response> { | |
| 204 | + | let gems = match self.gem_index(workspace, viewer).await? { | |
| 205 | + | Ok(gems) => gems, | |
| 206 | + | Err(refused) => return Ok(refused), | |
| 207 | + | }; | |
| 208 | + | let names: Vec<String> = gems.into_iter().map(|(p, _)| p.name).collect(); | |
| 209 | + | index_file(request, rubygems::names_file(&names), head) | |
| 210 | + | } | |
| 211 | + | ||
| 212 | + | async fn gem_info(&self, request: &Request, workspace: &str, name: &str, viewer: Option<&User>, head: bool) -> Result<Response> { | |
| 213 | + | let Some(package) = self.db.package(workspace, RUBYGEMS, name).await?.filter(|p| !p.hidden()) else { | |
| 214 | + | return self.gem_absent(workspace, viewer).await; | |
| 215 | + | }; | |
| 216 | + | if let Some(refusal) = self.gem_check(viewer, &package, Action::Pull).await? { | |
| 217 | + | return Ok(refusal); | |
| 218 | + | } | |
| 219 | + | let mut rows: Vec<VersionRow> = self.db.versions(&package.id, MAX_VERSIONS).await?.into_iter().filter(|v| !v.is_yanked()).collect(); | |
| 220 | + | if rows.is_empty() { | |
| 221 | + | return self.gem_absent(workspace, viewer).await; | |
| 222 | + | } | |
| 223 | + | rows.reverse(); | |
| 224 | + | index_file(request, rubygems::info(&rows.iter().map(line).collect::<Vec<_>>()), head) | |
| 225 | + | } | |
| 226 | + | ||
| 227 | + | /// A `.gem`, yanked ones too: a lockfile may still name them. | |
| 228 | + | async fn gem_download(&self, workspace: &str, stem: &str, viewer: Option<&User>, head: bool, ctx: &Context) -> Result<Response> { | |
| 229 | + | for (name, key) in rubygems::candidates(stem).into_iter().rev() { | |
| 230 | + | let Some(package) = self.db.package(workspace, RUBYGEMS, &name).await?.filter(|p| !p.hidden()) else { | |
| 231 | + | continue; | |
| 232 | + | }; | |
| 233 | + | if let Some(refusal) = self.gem_check(viewer, &package, Action::Pull).await? { | |
| 234 | + | return Ok(refusal); | |
| 235 | + | } | |
| 236 | + | let Some(row) = self.db.version_named(&package.id, &key).await? else { | |
| 237 | + | continue; | |
| 238 | + | }; | |
| 239 | + | let Some(digest) = Digest::parse(&row.digest) else { continue }; | |
| 240 | + | let Some(blob) = self.db.package_blob(&package.id, &digest).await? else { continue }; | |
| 241 | + | let headers = Headers::new(); | |
| 242 | + | headers.set("content-type", "application/octet-stream")?; | |
| 243 | + | headers.set("content-length", &blob.size.to_string())?; | |
| 244 | + | headers.set("cache-control", "max-age=31536000")?; | |
| 245 | + | if head { | |
| 246 | + | return Ok(Response::from_body(ResponseBody::Empty)?.with_headers(headers)); | |
| 247 | + | } | |
| 248 | + | let Some(got) = self.store.get(&blob.object_key, None).await? else { continue }; | |
| 249 | + | self.count_download(&package.id, ctx); | |
| 250 | + | return Ok(Response::from_body(got.body)?.with_headers(headers)); | |
| 251 | + | } | |
| 252 | + | self.gem_absent(workspace, viewer).await | |
| 253 | + | } | |
| 254 | + | ||
| 255 | + | /// The package a first push makes, linked to the repository the gem's | |
| 256 | + | /// `source_code_uri` or homepage names on g1t, or else one named like it. | |
| 257 | + | async fn gem_target(&self, workspace: &str, spec: &rubygems::Gemspec) -> Result<TargetOf> { | |
| 258 | + | let named: Vec<String> = [&spec.source_code_uri, &spec.homepage] | |
| 259 | + | .into_iter() | |
| 260 | + | .flatten() | |
| 261 | + | .filter_map(|url| npm::repository_of(&Value::String(url.clone()), &self.host)) | |
| 262 | + | .filter(|(owner, _)| owner == workspace) | |
| 263 | + | .map(|(_, repo)| repo) | |
| 264 | + | .collect(); | |
| 265 | + | let lower = spec.name.to_ascii_lowercase(); | |
| 266 | + | let dashed = lower.replace('_', "-"); | |
| 267 | + | let mut repo = None; | |
| 268 | + | for candidate in named.iter().map(String::as_str).chain([lower.as_str(), dashed.as_str()]) { | |
| 269 | + | if let Some(found) = self.repo_by_name(workspace, candidate).await? { | |
| 270 | + | repo = Some(found); | |
| 271 | + | break; | |
| 272 | + | } | |
| 273 | + | } | |
| 274 | + | Ok(TargetOf { workspace: workspace.to_owned(), repo: repo.map(|r| (r.id, r.name, r.is_private)), public: false }) | |
| 275 | + | } | |
| 276 | + | ||
| 277 | + | /// `gem push`: the `.gem` as the body. | |
| 278 | + | async fn gem_push(&self, request: &mut Request, workspace: &str, viewer: Option<&User>) -> Result<Response> { | |
| 279 | + | let declared = request.headers().get("content-length")?.and_then(|n| n.parse::<u64>().ok()); | |
| 280 | + | let too_large = || { | |
| 281 | + | let mb = self.max_request / 1_000_000; | |
| 282 | + | error(413, format!("A gem may be at most {mb} MB. See {DOCS}#size")) | |
| 283 | + | }; | |
| 284 | + | if declared.is_some_and(|n| n > self.max_request) { | |
| 285 | + | return too_large(); | |
| 286 | + | } | |
| 287 | + | let gem = request.bytes().await?; | |
| 288 | + | if gem.len() as u64 > self.max_request { | |
| 289 | + | return too_large(); | |
| 290 | + | } | |
| 291 | + | if viewer.is_none() { | |
| 292 | + | return error(401, format!("Push with a g1t access token: GEM_HOST_API_KEY=<token> gem push <file> --host https://g1t.sh/-/rubygems/{workspace}. Make one at {TOKENS}.")); | |
| 293 | + | } | |
| 294 | + | let spec = match rubygems::read_gem(&gem) { | |
| 295 | + | Ok(spec) => spec, | |
| 296 | + | Err(message) => return error(422, message), | |
| 297 | + | }; | |
| 298 | + | if !rubygems::valid_name(&spec.name) { | |
| 299 | + | return error(422, format!("{} is not a valid gem name: letters, digits, ., - and _, with a letter.", spec.name)); | |
| 300 | + | } | |
| 301 | + | if !rubygems::valid_version(&spec.version) { | |
| 302 | + | return error(422, format!("{} is not a version RubyGems reads.", spec.version)); | |
| 303 | + | } | |
| 304 | + | let key = rubygems::key(&spec.version, &spec.platform); | |
| 305 | + | ||
| 306 | + | // A name is taken whatever its case. | |
| 307 | + | let found = self.db.package_any_case(workspace, RUBYGEMS, &spec.name).await?; | |
| 308 | + | if let Some(found) = &found { | |
| 309 | + | if found.hidden() { | |
| 310 | + | return error(403, format!("The workspace {workspace} is deleted; nothing can be pushed to it.")); | |
| 311 | + | } | |
| 312 | + | if found.name != spec.name { | |
| 313 | + | return error(409, format!("The name {} is taken by the gem {}. Push it under that name.", spec.name, found.name)); | |
| 314 | + | } | |
| 315 | + | } else if self.db.workspace_hidden(workspace).await? { | |
| 316 | + | return error(403, format!("The workspace {workspace} is deleted; nothing can be pushed to it.")); | |
| 317 | + | } | |
| 318 | + | let target = match &found { | |
| 319 | + | Some(package) => TargetOf::package(package), | |
| 320 | + | None => self.gem_target(workspace, &spec).await?, | |
| 321 | + | }; | |
| 322 | + | let decision = access::decide(viewer, &target.view(), Action::Push); | |
| 323 | + | if !decision.allowed { | |
| 324 | + | let readable = found.is_none() || access::decide(viewer, &target.view(), Action::Pull).allowed; | |
| 325 | + | if !readable { | |
| 326 | + | return error(404, "Not found: no such gem, or you cannot see it."); | |
| 327 | + | } | |
| 328 | + | return error(403, decision.reason.unwrap_or_else(|| "Not allowed.".to_owned())); | |
| 329 | + | } | |
| 330 | + | let caller = Caller { actor: viewer.map(AuditActor::of) }; | |
| 331 | + | let package = match found { | |
| 332 | + | Some(package) => package, | |
| 333 | + | None => { | |
| 334 | + | self.db | |
| 335 | + | .create_package( | |
| 336 | + | &new_id("pkg", now_ms()), | |
| 337 | + | workspace, | |
| 338 | + | RUBYGEMS, | |
| 339 | + | &spec.name, | |
| 340 | + | target.repo.as_ref().map(|(id, repo, private)| (id.as_str(), repo.as_str(), *private)), | |
| 341 | + | caller.actor.as_ref().map_or("", |actor| actor.actor_id.as_str()), | |
| 342 | + | now_ms(), | |
| 343 | + | ) | |
| 344 | + | .await? | |
| 345 | + | } | |
| 346 | + | }; | |
| 347 | + | let existing = self.db.versions(&package.id, MAX_VERSIONS).await?; | |
| 348 | + | if existing.iter().any(|v| v.version == key) { | |
| 349 | + | return error(409, format!("{} ({key}) is already pushed, and a version is pushed once, yanked or not. Bump the version.", package.name)); | |
| 350 | + | } | |
| 351 | + | ||
| 352 | + | let digest = Digest::of(&gem); | |
| 353 | + | let size = gem.len() as u64; | |
| 354 | + | if let Some(refusal) = self.storage_refusal(&package, &[(digest.to_string(), size)]).await? { | |
| 355 | + | return error(403, refusal); | |
| 356 | + | } | |
| 357 | + | let now = now_ms(); | |
| 358 | + | let stored = match self.db.blob(&digest).await? { | |
| 359 | + | Some(blob) => self.store.head(&blob.object_key).await?.is_some(), | |
| 360 | + | None => false, | |
| 361 | + | }; | |
| 362 | + | if !stored { | |
| 363 | + | self.store.put(&digest.object_key(), gem).await?; | |
| 364 | + | } | |
| 365 | + | self.db.keep_blob(&package.id, &digest, size, Some("application/octet-stream"), &digest.object_key(), now).await?; | |
| 366 | + | self.db | |
| 367 | + | .publish( | |
| 368 | + | NewVersion { | |
| 369 | + | id: new_id("ver", now), | |
| 370 | + | package_id: package.id.clone(), | |
| 371 | + | version: key.clone(), | |
| 372 | + | digest: digest.to_string(), | |
| 373 | + | size, | |
| 374 | + | metadata: rubygems::stored(&spec).to_string(), | |
| 375 | + | subject: None, | |
| 376 | + | published_by: published_by(&caller), | |
| 377 | + | files: vec![NewFile { name: "gem".to_owned(), digest: digest.to_string(), size, media_type: Some("application/octet-stream".to_owned()) }], | |
| 378 | + | }, | |
| 379 | + | None, | |
| 380 | + | now, | |
| 381 | + | ) | |
| 382 | + | .await?; | |
| 383 | + | // The description the page shows: the highest stable version's. | |
| 384 | + | let highest = !rubygems::is_prerelease(&spec.version) | |
| 385 | + | && existing.iter().map(|v| v.meta()).filter_map(|m| m["number"].as_str().map(str::to_owned)).all(|other| { | |
| 386 | + | rubygems::is_prerelease(&other) || crate::db::newest_version(&format!("{other}\n{}", spec.version)).as_deref() == Some(spec.version.as_str()) | |
| 387 | + | }); | |
| 388 | + | if highest || existing.is_empty() { | |
| 389 | + | let description = spec.summary.as_deref().or(spec.description.as_deref()); | |
| 390 | + | self.db.set_readme(&package.id, None, description, now).await?; | |
| 391 | + | } | |
| 392 | + | self.db.measure(&package.workspace).await?; | |
| 393 | + | let event = PackageEvent { | |
| 394 | + | version: Some(key.clone()), | |
| 395 | + | digest: Some(digest.to_string()), | |
| 396 | + | size: Some(size), | |
| 397 | + | ..self.event_of(&package) | |
| 398 | + | }; | |
| 399 | + | self.announce("package.published", &package, event, &caller).await; | |
| 400 | + | self.audit(&caller, "package.publish", &package, Some(&format!("{workspace}/{}@{key}", package.name)), None).await; | |
| 401 | + | let mut response = Response::ok(format!("Successfully registered gem: {} ({key})", package.name))?; | |
| 402 | + | response.headers_mut().set("content-type", "text/plain; charset=utf-8")?; | |
| 403 | + | Ok(response) | |
| 404 | + | } | |
| 405 | + | ||
| 406 | + | /// `gem yank`: the version leaves the index, its file stays. | |
| 407 | + | async fn gem_yank(&self, request: &mut Request, url: &Url, workspace: &str, viewer: Option<&User>) -> Result<Response> { | |
| 408 | + | let body = request.text().await.unwrap_or_default(); | |
| 409 | + | let query = url.query().unwrap_or("").to_owned(); | |
| 410 | + | let value = |key: &str| rubygems::form_value(&body, key).or_else(|| rubygems::form_value(&query, key)); | |
| 411 | + | let (Some(name), Some(version)) = (value("gem_name"), value("version")) else { | |
| 412 | + | return error(400, "Name the gem and version: gem yank <gem> --version <version>."); | |
| 413 | + | }; | |
| 414 | + | let key = rubygems::key(&version, value("platform").as_deref().unwrap_or("ruby")); | |
| 415 | + | let Some(package) = self.db.package(workspace, RUBYGEMS, &name).await?.filter(|p| !p.hidden()) else { | |
| 416 | + | return self.gem_absent(workspace, viewer).await; | |
| 417 | + | }; | |
| 418 | + | if let Some(refusal) = self.gem_check(viewer, &package, Action::Push).await? { | |
| 419 | + | return Ok(refusal); | |
| 420 | + | } | |
| 421 | + | let Some(row) = self.db.version_named(&package.id, &key).await? else { | |
| 422 | + | return error(404, format!("{name} ({key}) is not there.")); | |
| 423 | + | }; | |
| 424 | + | if row.is_yanked() { | |
| 425 | + | return error(422, format!("{name} ({key}) is already yanked.")); | |
| 426 | + | } | |
| 427 | + | self.db.set_yanked(&row.id, true).await?; | |
| 428 | + | self.db.touch_package(&package.id, now_ms()).await?; | |
| 429 | + | let caller = Caller { actor: viewer.map(AuditActor::of) }; | |
| 430 | + | self.audit(&caller, "package.yank", &package, Some(&format!("{workspace}/{name}@{key}")), None).await; | |
| 431 | + | let mut response = Response::ok(format!("Successfully deleted gem: {name} ({key})"))?; | |
| 432 | + | response.headers_mut().set("content-type", "text/plain; charset=utf-8")?; | |
| 433 | + | Ok(response) | |
| 434 | + | } | |
| 435 | + | } |
| 1 | + | //! Just enough XML for package files: a Maven POM and a NuGet `.nuspec` | |
| 2 | + | //! are read into a tree of elements (names without their namespace | |
| 3 | + | //! prefix), and the metadata Maven reads is written with `escape`. | |
| 4 | + | //! Declarations, comments, doctypes and processing instructions are | |
| 5 | + | //! skipped; CDATA and the standard entities are read. | |
| 6 | + | ||
| 7 | + | #[derive(Clone, Debug, Default, PartialEq, Eq)] | |
| 8 | + | pub struct Element { | |
| 9 | + | /// The local name: `package` for `<ns:package>`. | |
| 10 | + | pub name: String, | |
| 11 | + | pub attributes: Vec<(String, String)>, | |
| 12 | + | pub children: Vec<Element>, | |
| 13 | + | /// The element's own text, its parts joined. | |
| 14 | + | pub text: String, | |
| 15 | + | } | |
| 16 | + | ||
| 17 | + | impl Element { | |
| 18 | + | /// The first child named `name`. | |
| 19 | + | pub fn child(&self, name: &str) -> Option<&Element> { | |
| 20 | + | self.children.iter().find(|c| c.name == name) | |
| 21 | + | } | |
| 22 | + | ||
| 23 | + | pub fn children_named<'a>(&'a self, name: &'a str) -> impl Iterator<Item = &'a Element> + 'a { | |
| 24 | + | self.children.iter().filter(move |c| c.name == name) | |
| 25 | + | } | |
| 26 | + | ||
| 27 | + | /// The trimmed text of the first child named `name`, when it has some. | |
| 28 | + | pub fn child_text(&self, name: &str) -> Option<String> { | |
| 29 | + | self.child(name).map(|c| c.text.trim().to_owned()).filter(|t| !t.is_empty()) | |
| 30 | + | } | |
| 31 | + | ||
| 32 | + | pub fn attribute(&self, name: &str) -> Option<&str> { | |
| 33 | + | self.attributes.iter().find(|(key, _)| key == name).map(|(_, value)| value.as_str()) | |
| 34 | + | } | |
| 35 | + | } | |
| 36 | + | ||
| 37 | + | fn local(name: &str) -> String { | |
| 38 | + | name.rsplit(':').next().unwrap_or(name).to_owned() | |
| 39 | + | } | |
| 40 | + | ||
| 41 | + | /// Text with its entities read: `&` is `&`, `A` is `A`. | |
| 42 | + | fn unescape(text: &str) -> String { | |
| 43 | + | let mut out = String::with_capacity(text.len()); | |
| 44 | + | let mut rest = text; | |
| 45 | + | while let Some(at) = rest.find('&') { | |
| 46 | + | out.push_str(&rest[..at]); | |
| 47 | + | rest = &rest[at..]; | |
| 48 | + | let Some(end) = rest.find(';').filter(|end| *end <= 12) else { | |
| 49 | + | out.push('&'); | |
| 50 | + | rest = &rest[1..]; | |
| 51 | + | continue; | |
| 52 | + | }; | |
| 53 | + | let entity = &rest[1..end]; | |
| 54 | + | let read = match entity { | |
| 55 | + | "amp" => Some('&'), | |
| 56 | + | "lt" => Some('<'), | |
| 57 | + | "gt" => Some('>'), | |
| 58 | + | "quot" => Some('"'), | |
| 59 | + | "apos" => Some('\''), | |
| 60 | + | _ => entity | |
| 61 | + | .strip_prefix("#x") | |
| 62 | + | .or_else(|| entity.strip_prefix("#X")) | |
| 63 | + | .and_then(|hex| u32::from_str_radix(hex, 16).ok()) | |
| 64 | + | .or_else(|| entity.strip_prefix('#').and_then(|n| n.parse().ok())) | |
| 65 | + | .and_then(char::from_u32), | |
| 66 | + | }; | |
| 67 | + | match read { | |
| 68 | + | Some(c) => { | |
| 69 | + | out.push(c); | |
| 70 | + | rest = &rest[end + 1..]; | |
| 71 | + | } | |
| 72 | + | None => { | |
| 73 | + | out.push('&'); | |
| 74 | + | rest = &rest[1..]; | |
| 75 | + | } | |
| 76 | + | } | |
| 77 | + | } | |
| 78 | + | out.push_str(rest); | |
| 79 | + | out | |
| 80 | + | } | |
| 81 | + | ||
| 82 | + | /// Text made safe to put between tags or in an attribute. | |
| 83 | + | pub fn escape(text: &str) -> String { | |
| 84 | + | text.replace('&', "&").replace('<', "<").replace('>', ">").replace('"', """) | |
| 85 | + | } | |
| 86 | + | ||
| 87 | + | fn attributes(text: &str) -> Result<Vec<(String, String)>, String> { | |
| 88 | + | let mut list = Vec::new(); | |
| 89 | + | let mut rest = text.trim(); | |
| 90 | + | while !rest.is_empty() { | |
| 91 | + | let eq = rest.find('=').ok_or("An attribute has no value.")?; | |
| 92 | + | let key = rest[..eq].trim(); | |
| 93 | + | let after = rest[eq + 1..].trim_start(); | |
| 94 | + | let quote = after.chars().next().filter(|c| *c == '"' || *c == '\'').ok_or("An attribute's value is not quoted.")?; | |
| 95 | + | let close = after[1..].find(quote).ok_or("An attribute's value is not closed.")? + 1; | |
| 96 | + | list.push((local(key), unescape(&after[1..close]))); | |
| 97 | + | rest = after[close + 1..].trim_start(); | |
| 98 | + | } | |
| 99 | + | Ok(list) | |
| 100 | + | } | |
| 101 | + | ||
| 102 | + | /// Reads a document into its root element. | |
| 103 | + | pub fn parse(text: &str) -> Result<Element, String> { | |
| 104 | + | let text = text.trim_start_matches('\u{feff}'); | |
| 105 | + | let mut stack: Vec<Element> = Vec::new(); | |
| 106 | + | let mut root = None; | |
| 107 | + | let mut rest = text; | |
| 108 | + | while !rest.is_empty() { | |
| 109 | + | let Some(open) = rest.find('<') else { | |
| 110 | + | if let Some(top) = stack.last_mut() { | |
| 111 | + | top.text.push_str(&unescape(rest)); | |
| 112 | + | } | |
| 113 | + | break; | |
| 114 | + | }; | |
| 115 | + | if open > 0 | |
| 116 | + | && let Some(top) = stack.last_mut() { | |
| 117 | + | top.text.push_str(&unescape(&rest[..open])); | |
| 118 | + | } | |
| 119 | + | rest = &rest[open..]; | |
| 120 | + | let skip = |rest: &str, end: &str| rest.find(end).map(|at| at + end.len()).ok_or_else(|| "The XML ends early.".to_owned()); | |
| 121 | + | if rest.starts_with("<?") { | |
| 122 | + | rest = &rest[skip(rest, "?>")?..]; | |
| 123 | + | } else if rest.starts_with("<!--") { | |
| 124 | + | rest = &rest[skip(rest, "-->")?..]; | |
| 125 | + | } else if let Some(cdata) = rest.strip_prefix("<![CDATA[") { | |
| 126 | + | let end = cdata.find("]]>").ok_or("The XML ends early.")?; | |
| 127 | + | if let Some(top) = stack.last_mut() { | |
| 128 | + | top.text.push_str(&cdata[..end]); | |
| 129 | + | } | |
| 130 | + | rest = &cdata[end + 3..]; | |
| 131 | + | } else if rest.starts_with("<!") { | |
| 132 | + | rest = &rest[skip(rest, ">")?..]; | |
| 133 | + | } else if let Some(closing) = rest.strip_prefix("</") { | |
| 134 | + | let end = closing.find('>').ok_or("The XML ends early.")?; | |
| 135 | + | let name = local(closing[..end].trim()); | |
| 136 | + | let element = stack.pop().ok_or("A tag is closed that was not opened.")?; | |
| 137 | + | if element.name != name { | |
| 138 | + | return Err(format!("<{}> is closed by </{name}>.", element.name)); | |
| 139 | + | } | |
| 140 | + | match stack.last_mut() { | |
| 141 | + | Some(parent) => parent.children.push(element), | |
| 142 | + | None => root = Some(element), | |
| 143 | + | } | |
| 144 | + | rest = &closing[end + 1..]; | |
| 145 | + | } else { | |
| 146 | + | // An opening tag; `>` inside a quoted attribute does not end it. | |
| 147 | + | let mut quote = None; | |
| 148 | + | let end = rest | |
| 149 | + | .char_indices() | |
| 150 | + | .find(|&(_, c)| { | |
| 151 | + | match quote { | |
| 152 | + | Some(q) if c == q => quote = None, | |
| 153 | + | None if c == '"' || c == '\'' => quote = Some(c), | |
| 154 | + | None if c == '>' => return true, | |
| 155 | + | _ => {} | |
| 156 | + | } | |
| 157 | + | false | |
| 158 | + | }) | |
| 159 | + | .map(|(at, _)| at) | |
| 160 | + | .ok_or("The XML ends early.")?; | |
| 161 | + | let inner = &rest[1..end]; | |
| 162 | + | let empty = inner.ends_with('/'); | |
| 163 | + | let inner = inner.trim_end_matches('/'); | |
| 164 | + | let (name, attrs) = inner.split_once(char::is_whitespace).unwrap_or((inner, "")); | |
| 165 | + | if name.is_empty() { | |
| 166 | + | return Err("A tag has no name.".to_owned()); | |
| 167 | + | } | |
| 168 | + | let element = Element { name: local(name), attributes: attributes(attrs)?, ..Element::default() }; | |
| 169 | + | if empty { | |
| 170 | + | match stack.last_mut() { | |
| 171 | + | Some(parent) => parent.children.push(element), | |
| 172 | + | None => root = Some(element), | |
| 173 | + | } | |
| 174 | + | } else { | |
| 175 | + | stack.push(element); | |
| 176 | + | } | |
| 177 | + | rest = &rest[end + 1..]; | |
| 178 | + | } | |
| 179 | + | if root.is_some() && stack.is_empty() { | |
| 180 | + | break; | |
| 181 | + | } | |
| 182 | + | } | |
| 183 | + | if !stack.is_empty() { | |
| 184 | + | return Err("The XML ends early.".to_owned()); | |
| 185 | + | } | |
| 186 | + | root.ok_or_else(|| "There is no XML in it.".to_owned()) | |
| 187 | + | } | |
| 188 | + | ||
| 189 | + | #[cfg(test)] | |
| 190 | + | mod tests { | |
| 191 | + | use super::*; | |
| 192 | + | ||
| 193 | + | #[test] | |
| 194 | + | fn documents_read_into_elements() { | |
| 195 | + | let doc = parse( | |
| 196 | + | r#"<?xml version="1.0" encoding="utf-8"?> | |
| 197 | + | <!-- a comment --> | |
| 198 | + | <package xmlns="http://schemas.microsoft.com/packaging/2013/05/nuspec.xsd"> | |
| 199 | + | <metadata minClientVersion="4.0"> | |
| 200 | + | <id>Acme.Web</id> | |
| 201 | + | <description><![CDATA[Fast & <small>]]></description> | |
| 202 | + | <authors>Ada & Bo</authors> | |
| 203 | + | <repository type="git" url="https://g1t.sh/acme/web.git" /> | |
| 204 | + | <dependencies> | |
| 205 | + | <group targetFramework="net8.0"><dependency id="Newtonsoft.Json" version="13.0.1" exclude="Build,Analyzers" /></group> | |
| 206 | + | </dependencies> | |
| 207 | + | <x:other xmlns:x="urn:x" note='a > b'>AB</x:other> | |
| 208 | + | </metadata> | |
| 209 | + | </package>"#, | |
| 210 | + | ) | |
| 211 | + | .unwrap(); | |
| 212 | + | assert_eq!(doc.name, "package"); | |
| 213 | + | let metadata = doc.child("metadata").unwrap(); | |
| 214 | + | assert_eq!(metadata.attribute("minClientVersion"), Some("4.0")); | |
| 215 | + | assert_eq!(metadata.child_text("id").as_deref(), Some("Acme.Web")); | |
| 216 | + | assert_eq!(metadata.child_text("description").as_deref(), Some("Fast & <small>")); | |
| 217 | + | assert_eq!(metadata.child_text("authors").as_deref(), Some("Ada & Bo")); | |
| 218 | + | assert_eq!(metadata.child("repository").unwrap().attribute("url"), Some("https://g1t.sh/acme/web.git")); | |
| 219 | + | let dep = metadata.child("dependencies").unwrap().child("group").unwrap().child("dependency").unwrap(); | |
| 220 | + | assert_eq!(dep.attribute("version"), Some("13.0.1")); | |
| 221 | + | let other = metadata.child("other").unwrap(); | |
| 222 | + | assert_eq!(other.text, "AB"); | |
| 223 | + | assert_eq!(other.attribute("note"), Some("a > b")); | |
| 224 | + | assert_eq!(metadata.child_text("missing"), None); | |
| 225 | + | } | |
| 226 | + | ||
| 227 | + | #[test] | |
| 228 | + | fn broken_documents_are_refused() { | |
| 229 | + | assert!(parse("<a><b></a>").is_err()); | |
| 230 | + | assert!(parse("<a>").is_err()); | |
| 231 | + | assert!(parse("just text").is_err()); | |
| 232 | + | assert!(parse("<a x=1/>").is_err()); | |
| 233 | + | assert_eq!(escape("a<b & \"c\""), "a<b & "c""); | |
| 234 | + | } | |
| 235 | + | } |
| 1 | + | //! Just enough YAML for a gem's `metadata.gz`: the `Gem::Specification` | |
| 2 | + | //! RubyGems writes with Psych. Block mappings and sequences (including a | |
| 3 | + | //! sequence at its key's indent, and `- - ">="` nested ones), plain and | |
| 4 | + | //! quoted scalars wrapped over lines, `|` and `>` block scalars, and empty | |
| 5 | + | //! flow collections. Tags (`!ruby/object:Gem::Version`) are dropped: the | |
| 6 | + | //! value is read as the plain mapping or scalar under them. | |
| 7 | + | ||
| 8 | + | use serde_json::{Map, Value}; | |
| 9 | + | ||
| 10 | + | #[derive(Clone, Debug)] | |
| 11 | + | struct Line { | |
| 12 | + | indent: usize, | |
| 13 | + | text: String, | |
| 14 | + | } | |
| 15 | + | ||
| 16 | + | /// Reads a document into JSON values: mappings as objects, sequences as | |
| 17 | + | /// arrays, scalars as strings, and empty values as null. | |
| 18 | + | pub fn parse(text: &str) -> Result<Value, String> { | |
| 19 | + | let mut lines = Vec::new(); | |
| 20 | + | for raw in text.lines() { | |
| 21 | + | let trimmed = raw.trim_end(); | |
| 22 | + | let indent = trimmed.len() - trimmed.trim_start().len(); | |
| 23 | + | let content = trimmed.trim_start(); | |
| 24 | + | if content.starts_with('#') || content == "..." { | |
| 25 | + | continue; | |
| 26 | + | } | |
| 27 | + | if indent == 0 && content.starts_with("---") { | |
| 28 | + | let after = strip_tag(content[3..].trim()); | |
| 29 | + | if !after.is_empty() { | |
| 30 | + | lines.push(Line { indent: 0, text: after.to_owned() }); | |
| 31 | + | } | |
| 32 | + | continue; | |
| 33 | + | } | |
| 34 | + | lines.push(Line { indent, text: content.to_owned() }); | |
| 35 | + | } | |
| 36 | + | let mut parser = Parser { lines, at: 0 }; | |
| 37 | + | parser.skip_blank(); | |
| 38 | + | if parser.at >= parser.lines.len() { | |
| 39 | + | return Ok(Value::Null); | |
| 40 | + | } | |
| 41 | + | let indent = parser.lines[parser.at].indent; | |
| 42 | + | parser.node(indent) | |
| 43 | + | } | |
| 44 | + | ||
| 45 | + | /// The text after a leading tag: `!ruby/object:Gem::Version` alone is empty. | |
| 46 | + | fn strip_tag(text: &str) -> &str { | |
| 47 | + | if text.starts_with('!') { | |
| 48 | + | text.split_once(' ').map_or("", |(_, rest)| rest.trim_start()) | |
| 49 | + | } else { | |
| 50 | + | text | |
| 51 | + | } | |
| 52 | + | } | |
| 53 | + | ||
| 54 | + | /// Where a mapping key ends: the `:` followed by a space or the end of the | |
| 55 | + | /// line, outside quotes. | |
| 56 | + | fn key_end(text: &str) -> Option<usize> { | |
| 57 | + | if text.starts_with('"') || text.starts_with('\'') { | |
| 58 | + | let quote = text.as_bytes()[0] as char; | |
| 59 | + | let close = text[1..].find(quote)? + 1; | |
| 60 | + | return text[close + 1..].starts_with(':').then_some(close + 1).filter(|at| { | |
| 61 | + | let after = &text[at + 1..]; | |
| 62 | + | after.is_empty() || after.starts_with(' ') | |
| 63 | + | }); | |
| 64 | + | } | |
| 65 | + | if text.starts_with('-') && (text.len() == 1 || text.as_bytes()[1] == b' ') { | |
| 66 | + | return None; | |
| 67 | + | } | |
| 68 | + | let bytes = text.as_bytes(); | |
| 69 | + | (0..bytes.len()).find(|&i| bytes[i] == b':' && (i + 1 == bytes.len() || bytes[i + 1] == b' ')) | |
| 70 | + | } | |
| 71 | + | ||
| 72 | + | fn unquote_key(key: &str) -> String { | |
| 73 | + | let key = key.trim(); | |
| 74 | + | match scalar(key) { | |
| 75 | + | Value::String(s) => s, | |
| 76 | + | _ => key.to_owned(), | |
| 77 | + | } | |
| 78 | + | } | |
| 79 | + | ||
| 80 | + | /// A scalar as written on one line (continuations already joined). | |
| 81 | + | fn scalar(text: &str) -> Value { | |
| 82 | + | let text = strip_tag(text.trim()); | |
| 83 | + | if text.is_empty() || text == "~" || text == "null" { | |
| 84 | + | return Value::Null; | |
| 85 | + | } | |
| 86 | + | if text == "[]" { | |
| 87 | + | return Value::Array(Vec::new()); | |
| 88 | + | } | |
| 89 | + | if text == "{}" { | |
| 90 | + | return Value::Object(Map::new()); | |
| 91 | + | } | |
| 92 | + | if let Some(inner) = text.strip_prefix('[').and_then(|t| t.strip_suffix(']')) { | |
| 93 | + | return Value::Array(inner.split(',').map(|item| scalar(item.trim())).collect()); | |
| 94 | + | } | |
| 95 | + | if let Some(inner) = text.strip_prefix('\'').and_then(|t| t.strip_suffix('\'')) { | |
| 96 | + | return Value::String(inner.replace("''", "'")); | |
| 97 | + | } | |
| 98 | + | if let Some(inner) = text.strip_prefix('"').and_then(|t| t.strip_suffix('"')) { | |
| 99 | + | let mut out = String::new(); | |
| 100 | + | let mut chars = inner.chars(); | |
| 101 | + | while let Some(c) = chars.next() { | |
| 102 | + | if c != '\\' { | |
| 103 | + | out.push(c); | |
| 104 | + | continue; | |
| 105 | + | } | |
| 106 | + | match chars.next() { | |
| 107 | + | Some('n') => out.push('\n'), | |
| 108 | + | Some('t') => out.push('\t'), | |
| 109 | + | Some('"') => out.push('"'), | |
| 110 | + | Some('\\') => out.push('\\'), | |
| 111 | + | Some('/') => out.push('/'), | |
| 112 | + | Some('0') => out.push('\0'), | |
| 113 | + | Some(' ') => out.push(' '), | |
| 114 | + | Some('x') => { | |
| 115 | + | let hex: String = chars.by_ref().take(2).collect(); | |
| 116 | + | out.extend(u32::from_str_radix(&hex, 16).ok().and_then(char::from_u32)); | |
| 117 | + | } | |
| 118 | + | Some('u') => { | |
| 119 | + | let hex: String = chars.by_ref().take(4).collect(); | |
| 120 | + | out.extend(u32::from_str_radix(&hex, 16).ok().and_then(char::from_u32)); | |
| 121 | + | } | |
| 122 | + | Some(other) => { | |
| 123 | + | out.push('\\'); | |
| 124 | + | out.push(other); | |
| 125 | + | } | |
| 126 | + | None => out.push('\\'), | |
| 127 | + | } | |
| 128 | + | } | |
| 129 | + | return Value::String(out); | |
| 130 | + | } | |
| 131 | + | Value::String(text.to_owned()) | |
| 132 | + | } | |
| 133 | + | ||
| 134 | + | struct Parser { | |
| 135 | + | lines: Vec<Line>, | |
| 136 | + | at: usize, | |
| 137 | + | } | |
| 138 | + | ||
| 139 | + | impl Parser { | |
| 140 | + | fn skip_blank(&mut self) { | |
| 141 | + | while self.at < self.lines.len() && self.lines[self.at].text.is_empty() { | |
| 142 | + | self.at += 1; | |
| 143 | + | } | |
| 144 | + | } | |
| 145 | + | ||
| 146 | + | fn peek(&mut self) -> Option<&Line> { | |
| 147 | + | self.skip_blank(); | |
| 148 | + | self.lines.get(self.at) | |
| 149 | + | } | |
| 150 | + | ||
| 151 | + | /// The node starting at the current line, which is at `indent`. | |
| 152 | + | fn node(&mut self, indent: usize) -> Result<Value, String> { | |
| 153 | + | let Some(line) = self.peek().cloned() else { | |
| 154 | + | return Ok(Value::Null); | |
| 155 | + | }; | |
| 156 | + | if line.text == "-" || line.text.starts_with("- ") { | |
| 157 | + | return self.sequence(line.indent); | |
| 158 | + | } | |
| 159 | + | if key_end(&line.text).is_some() { | |
| 160 | + | return self.mapping(line.indent); | |
| 161 | + | } | |
| 162 | + | self.at += 1; | |
| 163 | + | let text = self.continued(line.text.clone(), indent.saturating_sub(1)); | |
| 164 | + | Ok(scalar(&text)) | |
| 165 | + | } | |
| 166 | + | ||
| 167 | + | /// A scalar with the lines that continue it: deeper than `parent`, and | |
| 168 | + | /// not themselves a key or an item. | |
| 169 | + | fn continued(&mut self, mut text: String, parent: usize) -> String { | |
| 170 | + | let open_quote = |t: &str| { | |
| 171 | + | let t = t.trim(); | |
| 172 | + | (t.starts_with('"') && (t.len() == 1 || !t.ends_with('"') || t.ends_with("\\\""))) | |
| 173 | + | || (t.starts_with('\'') && (t.len() == 1 || !t.ends_with('\''))) | |
| 174 | + | }; | |
| 175 | + | loop { | |
| 176 | + | let quoted = open_quote(&text); | |
| 177 | + | let Some(next) = self.lines.get(self.at) else { break }; | |
| 178 | + | if next.text.is_empty() { | |
| 179 | + | if quoted { | |
| 180 | + | text.push('\n'); | |
| 181 | + | self.at += 1; | |
| 182 | + | continue; | |
| 183 | + | } | |
| 184 | + | break; | |
| 185 | + | } | |
| 186 | + | if next.indent <= parent || (!quoted && (key_end(&next.text).is_some() || next.text.starts_with("- "))) { | |
| 187 | + | break; | |
| 188 | + | } | |
| 189 | + | text.push(' '); | |
| 190 | + | text.push_str(&next.text); | |
| 191 | + | self.at += 1; | |
| 192 | + | } | |
| 193 | + | text | |
| 194 | + | } | |
| 195 | + | ||
| 196 | + | fn sequence(&mut self, indent: usize) -> Result<Value, String> { | |
| 197 | + | let mut items = Vec::new(); | |
| 198 | + | while let Some(line) = self.peek().cloned() { | |
| 199 | + | if line.indent != indent || !(line.text == "-" || line.text.starts_with("- ")) { | |
| 200 | + | break; | |
| 201 | + | } | |
| 202 | + | let rest = line.text[1..].trim_start(); | |
| 203 | + | let rest = strip_tag(rest).to_owned(); | |
| 204 | + | if rest.is_empty() { | |
| 205 | + | self.at += 1; | |
| 206 | + | match self.peek().cloned() { | |
| 207 | + | Some(next) if next.indent > indent => items.push(self.node(next.indent)?), | |
| 208 | + | _ => items.push(Value::Null), | |
| 209 | + | } | |
| 210 | + | continue; | |
| 211 | + | } | |
| 212 | + | // The item's content stands in for a line at its own column. | |
| 213 | + | let column = indent + (line.text.len() - rest.len()); | |
| 214 | + | self.lines[self.at] = Line { indent: column, text: rest }; | |
| 215 | + | items.push(self.node(column)?); | |
| 216 | + | } | |
| 217 | + | Ok(Value::Array(items)) | |
| 218 | + | } | |
| 219 | + | ||
| 220 | + | fn mapping(&mut self, indent: usize) -> Result<Value, String> { | |
| 221 | + | let mut map = Map::new(); | |
| 222 | + | while let Some(line) = self.peek().cloned() { | |
| 223 | + | if line.indent != indent { | |
| 224 | + | break; | |
| 225 | + | } | |
| 226 | + | let Some(end) = key_end(&line.text) else { break }; | |
| 227 | + | let key = unquote_key(&line.text[..end]); | |
| 228 | + | let rest = strip_tag(line.text[end + 1..].trim()).to_owned(); | |
| 229 | + | self.at += 1; | |
| 230 | + | let value = if rest.is_empty() { | |
| 231 | + | match self.peek().cloned() { | |
| 232 | + | Some(next) if next.indent > indent => self.node(next.indent)?, | |
| 233 | + | Some(next) if next.indent == indent && (next.text == "-" || next.text.starts_with("- ")) => self.sequence(indent)?, | |
| 234 | + | _ => Value::Null, | |
| 235 | + | } | |
| 236 | + | } else if let Some(style) = rest.strip_prefix('|').map(|s| (true, s)).or_else(|| rest.strip_prefix('>').map(|s| (false, s))) { | |
| 237 | + | self.block(indent, style.0, style.1) | |
| 238 | + | } else { | |
| 239 | + | let text = self.continued(rest, indent); | |
| 240 | + | scalar(&text) | |
| 241 | + | }; | |
| 242 | + | map.insert(key, value); | |
| 243 | + | } | |
| 244 | + | Ok(Value::Object(map)) | |
| 245 | + | } | |
| 246 | + | ||
| 247 | + | /// A `|` (literal) or `>` (folded) block scalar under a key at `indent`. | |
| 248 | + | fn block(&mut self, indent: usize, literal: bool, chomp: &str) -> Value { | |
| 249 | + | let mut lines: Vec<(usize, String)> = Vec::new(); | |
| 250 | + | while let Some(next) = self.lines.get(self.at) { | |
| 251 | + | if !next.text.is_empty() && next.indent <= indent { | |
| 252 | + | break; | |
| 253 | + | } | |
| 254 | + | lines.push((next.indent, next.text.clone())); | |
| 255 | + | self.at += 1; | |
| 256 | + | } | |
| 257 | + | while lines.last().is_some_and(|(_, t)| t.is_empty()) { | |
| 258 | + | lines.pop(); | |
| 259 | + | } | |
| 260 | + | let base = lines.iter().filter(|(_, t)| !t.is_empty()).map(|(i, _)| *i).min().unwrap_or(0); | |
| 261 | + | let shown: Vec<String> = lines | |
| 262 | + | .iter() | |
| 263 | + | .map(|(i, t)| if t.is_empty() { String::new() } else { format!("{}{t}", " ".repeat(i - base)) }) | |
| 264 | + | .collect(); | |
| 265 | + | let mut text = if literal { | |
| 266 | + | shown.join("\n") | |
| 267 | + | } else { | |
| 268 | + | let mut folded = String::new(); | |
| 269 | + | for (n, line) in shown.iter().enumerate() { | |
| 270 | + | if n > 0 { | |
| 271 | + | folded.push(if line.is_empty() || shown[n - 1].is_empty() { '\n' } else { ' ' }); | |
| 272 | + | } | |
| 273 | + | folded.push_str(line); | |
| 274 | + | } | |
| 275 | + | folded | |
| 276 | + | }; | |
| 277 | + | if !chomp.contains('-') && !text.is_empty() { | |
| 278 | + | text.push('\n'); | |
| 279 | + | } | |
| 280 | + | Value::String(text) | |
| 281 | + | } | |
| 282 | + | } | |
| 283 | + | ||
| 284 | + | #[cfg(test)] | |
| 285 | + | mod tests { | |
| 286 | + | use super::*; | |
| 287 | + | use serde_json::json; | |
| 288 | + | ||
| 289 | + | const SPEC: &str = r#"--- !ruby/object:Gem::Specification | |
| 290 | + | name: hello-world | |
| 291 | + | version: !ruby/object:Gem::Version | |
| 292 | + | version: 0.1.0 | |
| 293 | + | platform: ruby | |
| 294 | + | authors: | |
| 295 | + | - Ada Lovelace | |
| 296 | + | autorequire: | |
| 297 | + | bindir: exe | |
| 298 | + | cert_chain: [] | |
| 299 | + | date: 2026-10-06 00:00:00.000000000 Z | |
| 300 | + | dependencies: | |
| 301 | + | - !ruby/object:Gem::Dependency | |
| 302 | + | name: rack | |
| 303 | + | requirement: !ruby/object:Gem::Requirement | |
| 304 | + | requirements: | |
| 305 | + | - - ">=" | |
| 306 | + | - !ruby/object:Gem::Version | |
| 307 | + | version: '2.0' | |
| 308 | + | - - "<" | |
| 309 | + | - !ruby/object:Gem::Version | |
| 310 | + | version: '4' | |
| 311 | + | type: :runtime | |
| 312 | + | prerelease: false | |
| 313 | + | - !ruby/object:Gem::Dependency | |
| 314 | + | name: rspec | |
| 315 | + | requirement: !ruby/object:Gem::Requirement | |
| 316 | + | requirements: | |
| 317 | + | - - "~>" | |
| 318 | + | - !ruby/object:Gem::Version | |
| 319 | + | version: '3.12' | |
| 320 | + | type: :development | |
| 321 | + | description: |- | |
| 322 | + | Says hello. | |
| 323 | + | ||
| 324 | + | Then says it again. | |
| 325 | + | email: | |
| 326 | + | - ada@example.com | |
| 327 | + | homepage: https://g1t.sh/acme/hello-world | |
| 328 | + | licenses: | |
| 329 | + | - MIT | |
| 330 | + | metadata: | |
| 331 | + | source_code_uri: https://g1t.sh/acme/hello-world | |
| 332 | + | "quoted key": 'it''s' | |
| 333 | + | required_ruby_version: !ruby/object:Gem::Requirement | |
| 334 | + | requirements: | |
| 335 | + | - - ">=" | |
| 336 | + | - !ruby/object:Gem::Version | |
| 337 | + | version: 3.0.0 | |
| 338 | + | summary: A summary long enough that Psych wraps it onto a second line when it | |
| 339 | + | writes the specification | |
| 340 | + | test_files: [] | |
| 341 | + | "#; | |
| 342 | + | ||
| 343 | + | #[test] | |
| 344 | + | fn a_gemspec_reads_as_rubygems_wrote_it() { | |
| 345 | + | let spec = parse(SPEC).unwrap(); | |
| 346 | + | assert_eq!(spec["name"], "hello-world"); | |
| 347 | + | assert_eq!(spec["version"]["version"], "0.1.0"); | |
| 348 | + | assert_eq!(spec["platform"], "ruby"); | |
| 349 | + | assert_eq!(spec["authors"], json!(["Ada Lovelace"])); | |
| 350 | + | assert_eq!(spec["autorequire"], Value::Null); | |
| 351 | + | assert_eq!(spec["cert_chain"], json!([])); | |
| 352 | + | let deps = spec["dependencies"].as_array().unwrap(); | |
| 353 | + | assert_eq!(deps.len(), 2); | |
| 354 | + | assert_eq!(deps[0]["name"], "rack"); | |
| 355 | + | assert_eq!(deps[0]["type"], ":runtime"); | |
| 356 | + | assert_eq!(deps[0]["requirement"]["requirements"], json!([[">=", { "version": "2.0" }], ["<", { "version": "4" }]])); | |
| 357 | + | assert_eq!(deps[1]["type"], ":development"); | |
| 358 | + | assert_eq!(spec["description"], "Says hello.\n\nThen says it again."); | |
| 359 | + | assert_eq!(spec["metadata"]["source_code_uri"], "https://g1t.sh/acme/hello-world"); | |
| 360 | + | assert_eq!(spec["metadata"]["quoted key"], "it's"); | |
| 361 | + | assert_eq!(spec["required_ruby_version"]["requirements"][0][1]["version"], "3.0.0"); | |
| 362 | + | assert_eq!(spec["summary"], "A summary long enough that Psych wraps it onto a second line when it writes the specification"); | |
| 363 | + | assert_eq!(spec["test_files"], json!([])); | |
| 364 | + | } | |
| 365 | + | ||
| 366 | + | #[test] | |
| 367 | + | fn scalars_and_blocks() { | |
| 368 | + | let doc = parse("a: \"line\\nnext\"\nb: >\n folded\n text\nc: [x, 'y']\nd: ~\n").unwrap(); | |
| 369 | + | assert_eq!(doc["a"], "line\nnext"); | |
| 370 | + | assert_eq!(doc["b"], "folded text\n"); | |
| 371 | + | assert_eq!(doc["c"], json!(["x", "y"])); | |
| 372 | + | assert_eq!(doc["d"], Value::Null); | |
| 373 | + | assert_eq!(parse("").unwrap(), Value::Null); | |
| 374 | + | assert_eq!(parse("- a\n- b\n").unwrap(), json!(["a", "b"])); | |
| 375 | + | } | |
| 376 | + | } |