# Site Clones — Master Instructions

These rules apply to every job. Always read the job-level CLAUDE.md for per-job variables.

---

## Output Structure

```
[job-folder]/
├── index.html
├── assets/
│   ├── css/
│   ├── js/
│   ├── fonts/
│   ├── images/
│   └── media/
├── assets-new/        ← new media assets sent by the team (images, videos, etc.)
├── brief.pdf          ← optional — sent by the team for media asset placement context
├── ERRORS.md          ← created only if unresolved console errors exist
└── CLAUDE.md          ← per-job variables
```

---

## UI/UX Fidelity — Standing Rule

This applies throughout every phase. **Do NOT change any of the following:**
- Layout, spacing, or visual hierarchy
- CSS rules, class names, or inline styles
- Responsive breakpoints and media queries
- JavaScript behavior, animations, scroll effects, or interactions
- Font loading (Google Fonts `<link>`, Adobe Fonts, or self-hosted `@font-face`) — keep all intact
- Third-party UI libraries (Swiper, GSAP, Lottie, etc.) — keep all intact
- z-index, overflow, and positioning rules

The cloned page must be visually and functionally identical to the original on desktop, tablet, and mobile.

---

## Phase 1 — Rip & Clean

### Step 1 — Rip the Page

- Use `wget` to spider and download the page and all linked assets:
  ```bash
  wget --mirror --convert-links --adjust-extension --page-requisites --no-parent -P [job-folder] [URL]
  ```
- Move the downloaded HTML to `index.html` at the job root
- Move all other files (CSS, JS, images, fonts, media) into `assets/` with subfolders as above
- Rewrite all asset references in `index.html`, CSS, and JS to reflect new paths under `assets/`
- Verify no references still point to the original domain

**No external CDN — download everything locally.** After moving files, scan `index.html` (and any CSS/JS) for remaining external image/font/media URLs (e.g. `cloudfront.net`, `imgix.net`, CDN subdomains). Download each file into the appropriate `assets/` subfolder with `wget -q [url]` and replace the URL in the HTML/CSS with the local `assets/…` path. The finished clone must not depend on any external CDN for visual assets.

### Step 2 — Strip Tracking & Analytics

Remove all of the following from ALL `.html` and `.js` files:

**Script tags and inline code for:**
- Google Analytics / GTM (ga.js, analytics.js, gtag.js, GTM, Analyzely)
- Meta Pixel (fbevents.js, facebook.com/tr)
- Snap Pixel (sc-static.net/scevent.min.js, `snaptr(`)
- Taboola pixel (archive-digger.com/assets/script_tag.js, taboolaId)
- Hotjar, Intercom, Segment, Mixpanel, Heap, FullStory, Clarity, Amplitude
- TripleWhale / Triple Pixel (config-security.com)
- Elevate A/B testing (elevateab.app.txt, ds0wlyksfn0sb.cloudfront.net)
- Shoplift A/B testing (`<!-- Start of Shoplift scripts -->` ... `<!-- End of Shoplift scripts -->`)
- Klaviyo, ReCharge, Okendo (tracking calls only — keep Okendo CSS/widget HTML for visual fidelity)
- Any `<script>` tags loading from third-party analytics or tracking domains
- Any `fetch()`, `XHR`, or `navigator.sendBeacon()` calls to analytics endpoints
- Cookie consent banners tied to analytics (OneTrust, Cookiebot, etc.)

**Shopify platform scripts to remove (non-functional locally):**
- `window.Shopify.SignInWithShop?.initShopCartSync` — triggers `/cart.js` fetches
- `window.Shopify.featureAssets` — loads shop-cart-sync, checkout-modal, etc.
- Kaching Bundles app block (`shopify://apps/kaching-bundles`) and its `kaching-loader.js` script tag
- Cloudflare email-decode (`/cdn-cgi/scripts/.../email-decode.min.js`)
- Any other `shopify://apps/` embed blocks that are purely functional (cart, upsell, checkout)

**Also remove:**
- `<noscript>` pixel tags (Meta, GTM, etc.)
- Hidden tracking `<img>` tags (1x1 pixels)
- Any `data-*` attributes that reference tracking IDs

---

## Phase 2 — Apply Team Changes

*Run this phase only when the team provides changes. All steps are optional — apply only what is provided and skip the rest.*

### Step 3 — Read Brief (if present)

- **Text changes** arrive as plain text in the thread — this is preferred. Apply them directly with no PDF needed.
- **`brief.pdf`** is used when the team needs to show visual asset placement (which new image/video goes in which section). If present, read it to extract the asset swap map. Also check for any text changes it may contain.
- If the PDF shows slides side-by-side (original vs new), treat left as original and right as replacement.
- If neither is present, skip to Phase 3.

### Step 4 — Brand, Product & Text Changes

**Brand/product sweep** — only if Original brand and Our brand are present in the per-job CLAUDE.md (not blank or `(TBD)`). Replace across ALL files (HTML, JS, CSS, meta tags, JSON-LD, OG tags):
- `[ORIGINAL_BRAND]` → `[OUR_BRAND]`
- `[ORIGINAL_PRODUCT]` → `[OUR_PRODUCT]`

If branding is missing, skip the brand/product sweep and leave a note in `ERRORS.md` that branding is pending.

**Specific text changes from thread or brief** — apply each one precisely:
- Match the exact string in the HTML (use the original text from the page as the search key)
- Replace with the exact new copy provided
- Preserve surrounding HTML tags, classes, and formatting — only the text content changes
- If a text change affects a button label, also check if the CTA link needs updating (cross-reference Step 5)

**Also update:**
- `<title>` tag
- `<meta name="description">`
- Open Graph tags (`og:title`, `og:description`, `og:site_name`)
- Twitter card tags
- JSON-LD structured data (name, brand fields)
- Alt text on brand images/logos

### Step 5 — CTA Link Replacement

Read CTA URL from the job-level CLAUDE.md. If it is missing or `(TBD)`, skip this step entirely and leave a note in `ERRORS.md` that the CTA URL is pending.

- Replace ALL primary CTA button `href` values with `[OUR_CTA_URL]`
- Check for JS-driven navigation (onClick handlers, router pushes, window.location assignments) and replace those too
- Check for form `action` attributes pointing to original domains
- Do NOT change navigation links, footer links, or non-CTA internal links

**Goto() / indirect CTA pattern** — many pages use a `Goto()` function that builds a redirect URL and a link-rewriting script that sets all non-footer `<a>` tags to `href='#'` with a click listener calling `Goto()`. Remove this pattern entirely and replace with direct links:
1. Delete the `Goto()` function, `getQueryString()`, and `GetRequest()` helpers
2. In the link-rewriting loop, replace the `else` branch (the one that was calling `Goto()`) with: `currentA.href = '[OUR_CTA_URL]'; currentA.target = '_blank';`
3. Leave `privacy-link`, `articles_links`, and `image-misalignment-link` branches unchanged

### Step 6 — Media Asset Replacement

Use the asset swap map from `brief.pdf` (Step 3) and/or the table in the per-job CLAUDE.md.

- New assets are in `assets-new/` (sent by the team — any format: JPG, PNG, MP4, etc.)
- For each swap: copy the new asset into the correct `assets/` subfolder and rename it to match the original filename so no HTML/CSS references break — OR update all references if renaming isn't appropriate
- If the brief shows the new asset visually placed in a specific section of the page, match it to the image in that section in the HTML
- Preserve `width`, `height`, and `alt` attributes on replaced `<img>` tags
- If no new assets are provided, skip this step

---

## Phase 3 — Final QA

### Step 7 — Console Error Audit

1. **Always serve via HTTP — never open `index.html` directly as a `file://` URL.** Fonts and cross-origin scripts will be blocked by CORS on `file://`. Use:
   ```bash
   python3 -m http.server 8080
   ```
   Then open `http://localhost:8080` in the browser.
2. Open DevTools → Console and Network tabs. **Scroll the full page** to trigger lazy-loaded assets and scroll-event JS.
3. Fix ALL of the following:
   - Broken asset paths (404s for CSS, JS, fonts, images)
   - Mixed content warnings (http vs https)
   - Missing files referenced in CSS (`url()` paths)
   - CORS errors on self-hosted assets
   - JS runtime errors fired on scroll/resize — add null guards (`if (!el) return;`) inside event handler functions that query DOM elements by class/ID
   - Image src paths with `&width=N` instead of `?width=N` — replace `&` with `?` (caused by CDN URL construction when the base path already contained a query param)
   - JS errors caused by removed tracking scripts (wrap in try/catch or remove dependent code cleanly)
4. For errors that cannot be resolved (e.g. Okendo reviews API, Shopify cart/checkout), document them in `ERRORS.md`
5. Goal: zero critical errors. Non-critical warnings are acceptable if documented.

---

## Next.js Sites — Special Handling

Next.js pages cannot be cloned like a regular static site. The framework requires specific inline scripts and JS chunks to hydrate — removing or misrouting them produces a blank page or 404.

### How to detect Next.js

Page HTML contains any of: `__NEXT_DATA__`, `__next_f`, `/_next/static/`, `next/dist`, or `_next/` in script `src` attributes.

### Step 1 — wget adjustments for Next.js

The standard `wget --mirror` command mishandles Next.js chunk paths because they live under `_next/` and are content-hashed. Use this instead:

```bash
wget --mirror --convert-links --adjust-extension --page-requisites --no-parent \
  --reject "*.map" -P [job-folder] [URL]
```

After downloading, verify that `_next/static/chunks/` exists and contains `.js` files. If the directory is missing or empty, manually download the missing chunks by inspecting the Network tab (filter by JS, copy the chunk URLs, wget each one).

### Step 2 — Scripts you must NEVER remove on Next.js pages

Even though the cleaning step removes inline `<script>` blocks, these are **off-limits** on Next.js pages:

| Pattern | Why it must stay |
|---|---|
| `<script id="__NEXT_DATA__" type="application/json">` | Contains initial props, build ID, and route data. Removing it kills hydration entirely — blank page. |
| `<script src="/_next/static/...">` | Framework and page chunks. Removing any of these breaks the React app. |
| Any inline `<script>` referencing `__NEXT_F`, `__NEXT_P`, `self.__next_f`, `__webpack_require__`, or `webpackChunk` | Next.js runtime internals — not tracking. |
| `<script id="__NEXT_FONT_MANIFEST__">` or similar Next.js manifest scripts | Font and route manifests needed by the router. |

### Step 3 — Path rewriting for `_next/` assets

When rewriting asset paths after moving files, preserve the `_next/` directory name exactly. Do **not** move `_next/` contents into `assets/` — keep them at:

```
[job-folder]/_next/static/chunks/
[job-folder]/_next/static/css/
[job-folder]/_next/static/media/
```

References to `/_next/` in HTML and JS should become `_next/` (relative, no leading slash) when served from the job root.

**After copying files, fix subpath references in both `index.html` and all JS chunks.** Next.js sites hosted under a subpath (e.g. `/breezebox/`) embed that subpath as hardcoded strings in two places that `wget --convert-links` does NOT fix:

1. **Turbopack chunk loader** — one of the `_next/static/chunks/turbopack-*.js` files contains a string like `"/[subpath]/_next/"` that it uses as the base URL for dynamically loading all other chunks. Find and replace it:
   ```bash
   # Find which chunk has it:
   grep -rl '/[subpath]/_next/' _next/static/chunks/
   # Fix it (replace with just "_next/"):
   sed -i 's|"/[subpath]/_next/"|"_next/"|g' _next/static/chunks/turbopack-*.js
   ```
   If this string is wrong, the page loads blank — React chunks never load.

2. **Page component chunk** — one of the `_next/static/chunks/*.js` files contains a variable assignment like `e="/[subpath]"` (or `t="/[subpath]"`, `r="/[subpath]"`) that is used as the base URL prefix for all CSS, images, JS, and other static assets via template literals (`${e}/css/...`, `${e}/images/...`). Blank it out:
   ```bash
   # Find which chunk has it:
   grep -rl 'e="/[subpath]"' _next/static/chunks/
   # Fix it:
   sed -i 's|e="/[subpath]"|e=""|g' _next/static/chunks/[page-chunk].js
   ```
   With `e=""`, all asset paths become root-relative (e.g. `/css/pre/bootstrap.min.css`). You must then download those assets from the origin and place them at matching paths under the job root.

**Also apply the same `/[subpath]/` → `/` replacement in `index.html`** for any occurrences in inline `__next_f` script blocks (these also bypass `--convert-links`):
```python
content = content.replace('/[subpath]/_next/', '_next/')
content = content.replace('/[subpath]/', '/')
```

**Downloading assets referenced via the page chunk base variable:**
After blanking `e`, grep the page chunk for all `${e}/...` patterns to build a complete list of CSS, JS, images, fonts, and media the page loads:
```bash
grep -oh '\${e}/[^"'\''`,) ]*' _next/static/chunks/[page-chunk].js | sort -u
```
Download each file from the origin into the matching local path. Common structure: `css/[slug]/`, `images/[slug]/`, `js/[slug]/`. If font files (`.woff2`, `.woff`, `.ttf`) 404 on the origin, download them from cdnjs using the version found in the CSS file header.

### Step 4 — API route failures (expected, non-fixable)

Next.js pages often call `/api/...` endpoints on load to fetch dynamic content. These will 404 locally — that is expected and cannot be fixed without a running server. Document them in `ERRORS.md`.

### Diagnostic checklist — blank page / 404

Open DevTools before assuming a tracking script was wrongly removed:

1. **Elements tab** — is `<script id="__NEXT_DATA__">` present? If missing, the cleaning step removed it. Restore it from the original downloaded file.
2. **Console tab** — look for:
   - `Hydration failed` → `__NEXT_DATA__` is missing or corrupted
   - `Cannot read properties of undefined (reading 'call')` → a webpack chunk is missing (404 on a `_next/static/chunks/*.js` file)
   - `ChunkLoadError` → same as above — a JS chunk didn't load
3. **Network tab (filter: JS)** — are any `_next/static/chunks/*.js` files returning 404? If so, re-download the missing chunks and place them at the correct relative path.
4. **Network tab (filter: Fetch/XHR)** — `/api/` calls returning 404 are expected and non-critical. Document them in `ERRORS.md`.
5. **Page shows text but CSS/images 404** — the page chunk base variable (`e="/[subpath]"`) was not blanked. Fix it as described in Step 3.

---

## GemPages Pages — Local Asset Path Fix

If the source page is built with GemPages (Shopify app), the downloaded `assets/js/gp-lazyload.js` contains a broken `A()` function that prepends `https://` to relative paths, turning `assets/images/foo.jpg` into `https://assets/images/foo.jpg` (treating `assets` as a hostname). This breaks all locally-hosted images and media.

After downloading, find and replace this exact function in `gp-lazyload.js`:

**Find:**
```
let A=(e,t)=>{try{e.startsWith("http://")||e.startsWith("https://")||(e="https://"+e);let n=new URL(e);return n.searchParams.delete(t),n.toString()}catch(t){return console.error("Error occurred while removing query by key:",t),e}}
```

**Replace with:**
```
let A=(e,t)=>{try{var isRel=!e.startsWith("http://")&&!e.startsWith("https://");var base=isRel?window.location.origin+"/"+e.replace(/^\//,""):e;let n=new URL(base);n.searchParams.delete(t);return isRel?n.pathname.replace(/^\//,"")+n.search+n.hash:n.toString()}catch(t){return console.error("Error occurred while removing query by key:",t),e}}
```

**How to detect GemPages:** page HTML contains `gp_lazyload`, `gps-link`, `base-src`, or `gp-global.js`.

---

## Final Checklist Before Done

- [ ] `index.html` exists at job root
- [ ] All assets under `assets/` with no broken paths
- [ ] No tracking/analytics scripts remaining
- [ ] Brand and product names fully replaced — OR noted as pending in `ERRORS.md`
- [ ] All text changes from thread or brief applied (if provided)
- [ ] CTA links replaced — OR noted as pending in `ERRORS.md`
- [ ] New media assets in place and matched to correct page sections (if provided)
- [ ] Page looks identical on desktop, tablet, mobile
- [ ] Zero critical console errors
- [ ] `ERRORS.md` created if any unresolved issues or pending items
