2 核 7G 服务器磁盘 93%、CPU 160%?三步排查 Docker 资源黑洞
在为客户维护一台 2 核 7G 的云服务器时,发现磁盘使用率 93%(仅剩 3G),Airflow scheduler 持续占用 160% CPU。以下是完整的诊断和修复过程。
在为客户维护一台 2 核 7G 的云服务器时,发现磁盘使用率 93%(仅剩 3G),Airflow scheduler 持续占用 160% CPU。以下是完整的诊断和修复过程。
在为客户构建 Node.js 后端时,登录接口返回 401,排查发现 JWT token 全部验证失败。日志显示 JWT_SECRET 的值是字符串 "undefined" 而非真正的密钥。记录根因与解法。
在为客户构建数据采集工具时遇到此问题,记录从 Puppeteer 到 Electron 的完整排查过程。
Puppeteer stealth plugin 无法绕过高级验证码反爬系统。切换到 Chrome CDP 远程调试方案后,又遇到 WSL2 网络隔离和 Chrome 单实例限制两个坑。最终用 Electron BrowserWindow 加载目标网站,用户手动登录后通过 session.cookies.get() 自动提取 Cookie,彻底解决了反爬检测和跨平台问题。
使用 puppeteer-extra + puppeteer-extra-plugin-stealth 自动登录目标网站,浏览器启动后立即触发验证码拦截。即使通过了验证码,后续页面也会再次检测到自动化环境并强制退出。
import puppeteer from 'puppeteer-extra';
import StealthPlugin from 'puppeteer-extra-plugin-stealth';
puppeteer.use(StealthPlugin());
const browser = await puppeteer.launch({
headless: false,
args: [
'--disable-blink-features=AutomationControlled',
'--no-sandbox',
'--disable-infobars',
],
});
const page = await browser.newPage();
await page.goto('https://target-site.com/login');
// 验证码系统检测到自动化环境,页面被拦截
高级反爬验证码系统不只检查 navigator.webdriver 等基础指纹。它通过多个维度判断自动化环境:Chromium 编译特征、Canvas/WebGL 渲染差异、鼠标轨迹模式、甚至 DevTools Protocol 调用栈。stealth plugin 能修复已知的指纹泄露点,但无法消除 Puppeteer Chromium 与正常 Chrome 的底层差异。
经过 8 轮排查(移除超时、监听断开事件、排查环境差异、stealth plugin 补全等),确认无论怎么配置都无法绕过。
放弃 Puppeteer 自动登录,改用 Chrome DevTools Protocol (CDP) 连接用户真实浏览器提取 Cookie。
在 WSL2 中运行 Node.js 脚本,通过 chrome-remote-interface 连接 Windows Chrome 的调试端口,连接超时:
import CDP from 'chrome-remote-interface';
const client = await CDP({
host: 'localhost',
port: 9222,
});
// Error: connect ECONNREFUSED 127.0.0.1:9222
Windows 端启动 Chrome 的命令:
chrome.exe --remote-debugging-port=9222
在 Windows PowerShell 中 curl localhost:9222/json 正常返回,但 WSL2 内无法连接。
Chrome --remote-debugging-port=9222 默认绑定 127.0.0.1,即 Windows 的本地回环地址。WSL2 和 Windows 有各自独立的网络栈——WSL2 内的 localhost 指向 Linux 的回环地址,不是 Windows 的。所以从 WSL2 访问 localhost:9222 实际访问的是 Linux 的 9222 端口,而非 Windows Chrome。
启动 Chrome 时加 --remote-debugging-address=0.0.0.0,让 CDP 监听所有网卡:
chrome.exe --remote-debugging-port=9222 --remote-debugging-address=0.0.0.0
或者用 Windows netsh 配置端口转发:
netsh interface portproxy add v4tov4 listenport=9222 listenaddress=0.0.0.0 connectport=9222 connectaddress=127.0.0.1
--remote-debugging-address=0.0.0.0 会将 CDP 端口暴露给局域网,存在安全风险。仅在内网开发环境使用,生产环境务必配合防火墙规则限制访问来源。
Chrome 已经在运行,带 --remote-debugging-port=9222 参数重新启动,参数被静默忽略。Chrome 只是在已有窗口中打开新标签页,CDP 端口没有开启:
# Chrome 已在运行
chrome.exe --remote-debugging-port=9222
# 没有报错,但 9222 端口并未监听
Chrome 设计为单实例应用。检测到已有 Chrome 进程时,新启动的 Chrome 会将启动参数中的 URL 转发给已有进程,然后自行退出。--remote-debugging-port 等参数只在进程首次创建时生效,已有进程不会动态加载。
先关闭所有 Chrome 进程,再带参数重启:
# Windows
taskkill /F /IM chrome.exe
chrome.exe --remote-debugging-port=9222
# macOS
pkill -f "Google Chrome"
/Applications/Google\ Chrome.app/Contents/MacOS/Google\ Chrome --remote-debugging-port=9222
也可以用 --user-data-dir 指定独立配置目录,避免与日常使用的 Chrome 冲突:
chrome.exe --remote-debugging-port=9222 --user-data-dir="C:\chrome-debug-profile"
三个场景的踩坑说明了一个事实:依赖外部 Chrome 进程做 Cookie 提取,在 WSL2 + 反爬检测 + 进程管理的组合下太脆弱。最终方案是用 Electron 的 BrowserWindow 替代外部 Chrome。
打开登录窗口并轮询 Cookie:
import { BrowserWindow } from 'electron';
let cookieWindow = null;
function openLoginWindow() {
cookieWindow = new BrowserWindow({
width: 1000,
height: 700,
title: '登录目标平台',
webPreferences: {
// 关键:使用独立 session,不影响主窗口
partition: 'cookie-login',
contextIsolation: true,
nodeIntegration: false,
},
});
cookieWindow.loadURL('https://target-site.com/login');
// 轮询检测目标 Cookie
const interval = setInterval(async () => {
const cookies = await cookieWindow.webContents.session.cookies.get({
domain: '.target-site.com',
});
const sessionCookie = cookies.find((c) => c.name === 'session_token');
if (sessionCookie) {
clearInterval(interval);
// 拼接完整 Cookie 字符串
const cookieStr = cookies
.map((c) => c.name + '=' + c.value)
.join('; ');
// 写入环境变量
process.env.SESSION_COOKIE = cookieStr;
updateEnvFile('SESSION_COOKIE', cookieStr);
cookieWindow.close();
}
}, 2000);
cookieWindow.on('closed', () => {
clearInterval(interval);
cookieWindow = null;
});
}
preload 脚本暴露 IPC 接口:
import { contextBridge, ipcRenderer } from 'electron';
contextBridge.exposeInMainWorld('electronAPI', {
extractCookie: () => ipcRenderer.invoke('extract-cookie'),
onCookieExtracted: (callback) =>
ipcRenderer.on('cookie-extracted', (_, data) => callback(data)),
});
渲染进程调用:
// 检测是否在 Electron 环境
if (window.electronAPI) {
document.getElementById('extractBtn').addEventListener('click', () => {
window.electronAPI.extractCookie();
});
window.electronAPI.onCookieExtracted((data) => {
console.log('Cookie 已提取:', data);
});
}
partition: 'cookie-login' 创建隔离 session,登录窗口的 Cookie 不会污染主窗口。如果需要共享登录状态,去掉 partition 参数或使用相同 partition 名while + await 替代 setInterval,那会阻塞渲染进程session.cookies.get() 只能获取当前 session 的 Cookie,无法跨 partition 读取stealth plugin 只能绕过基础指纹检测(如 navigator.webdriver),高级反爬系统通过浏览器行为和底层 API 特征识别自动化环境。检测维度包括 Chromium 编译特征、Canvas 渲染差异、鼠标轨迹模式等,stealth plugin 无法完全覆盖。改用 Electron BrowserWindow 加载目标网站,用户手动登录后通过 session.cookies.get() 提取 Cookie 即可。
Chrome 采用单进程架构,已有 Chrome 进程运行时 --remote-debugging-port 参数会被静默忽略。新启动的 Chrome 只会在现有实例中打开新标签页,调试端口不会开启。需要先用 taskkill /F /IM chrome.exe(Windows)或 pkill -f "Google Chrome"(macOS)关闭所有 Chrome 进程,再带参数重新启动。也可以用 --user-data-dir 指定独立配置目录避免冲突。
Chrome 的 --remote-debugging-port 默认绑定 Windows 的 127.0.0.1,而 WSL2 有独立的网络栈,localhost 指向 Linux 回环地址而非 Windows。两种解法:一是启动 Chrome 时加 --remote-debugging-address=0.0.0.0 让 CDP 监听所有网卡;二是在 Windows 上用 netsh interface portproxy 配置端口转发,将 WSL2 的请求转发到 Windows 的 CDP 端口。
需要数据采集或自动化工具开发?
联系我们在将 Hono 后端部署到 Vercel Serverless Function 时,遇到三个层层递进的问题:esbuild 打包格式报错、原生依赖无法打包、以及最隐蔽的 catch-all 路由无法匹配多级路径。记录完整排查过程和最终方案。
--format=cjs 输出,原生依赖 --external 排除api/[[...path]].ts 不可靠,改用 api/index.ts + vercel.json rewritebaseURL 必须完整 URL,用 window.location.origin 拼接api/[[...path]].ts 直接 import Hono 的 app.ts,Vercel 用内置 TypeScript 5.9.3(nodenext 模式)编译,报出一串错误:
Relative import paths need explicit file extensions
Cannot find name 'process'
Module '"@libsql/client"' declares 'Client' locally, but it is not exported
根因是 Vercel 的内置 TS 编译器对 nodenext 模块解析要求严格,而 Hono + Turso 客户端库的导出方式不兼容。
解决方案:用 esbuild 预打包,让 Vercel 直接执行编译产物。
esbuild server/src/app.ts \
--bundle --platform=node --format=cjs \
--outfile=dist/_server.cjs \
--external:@libsql/client
api/[[...path]].ts 只需一行引用打包产物:
import app from '../../dist/_server.cjs';
export default app;
--format=esm 全量打包时,dotenv 内部 require("path") 报 Dynamic require is not supported。这是 ESM 规范限制,CJS 没有这个问题。
如果你在 Node.js ESM 动态 import 中也遇到过模块解析问题,根因是一样的——ESM 对模块格式要求严格,CJS 更宽容。
@libsql/client 包含原生二进制(@libsql/linux-x64-gnu),esbuild 打包后运行时找不到模块。解决方案:
--external:@libsql/client 排除打包package.json 声明 @libsql/client 为 dependency,让 Vercel 安装到 node_modulesesbuild 打包解决后,一级路由 /api/health、/api/tasks 全部正常。但所有多级路径(如 /api/auth/sign-in/email)返回 Vercel 层 404:
HTTP/2 404
x-vercel-error: NOT_FOUND
content-type: text/plain; charset=utf-8
通过对比 response headers 定位断裂点:
| 路径 | 到达 Hono? | 关键特征 |
|---|---|---|
/api/health | ✅ | x-vercel-cache: MISS、有 CORS header |
/api/gate | ✅ | 有 CORS header |
/api/gate/session | ❌ | x-vercel-error: NOT_FOUND、无 CORS |
一级路径到达函数,二级路径被拦截。 与路径名无关(不含 auth 的路径也 404),与 Vercel 内部路由规则无关。根因是 api/[[...path]].ts 的 catch-all pattern 只能匹配单级路径。
放弃 catch-all,改用 api/index.ts + vercel.json rewrite:
{
"rewrites": [
{ "source": "/api/(.*)", "destination": "/api/index" },
{ "source": "/((?!api/).*)", "destination": "/index.html" }
]
}
api/[[...path]].ts 重命名为 api/index.ts(内容不变)。所有 /api/* 请求由 rewrite 统一转发到 /api/index 函数,Hono 内部自行路由。
部署后验证:
curl -sI https://example.com/api/gate/session
# HTTP/2 404
# access-control-allow-credentials: true ← 到达 Hono
# x-vercel-cache: MISS ← 不再是 Vercel 层 404
三条路径 /api/gate、/api/gate/、/api/gate/session 全部到达 Hono。POST 请求也正常:
curl -X POST https://example.com/api/gate/sign-in/email \
-H "Content-Type: application/json" \
-d '{"email":"test@test.com","password":"123456"}'
# {"message":"Invalid email or password","code":"INVALID_EMAIL_OR_PASSWORD"}
Auth 路由完全打通。
/api/(.*) 必须在 SPA fallback 之前路由打通后,前端页面报错:
BetterAuthError: Invalid base URL: /api/gate
Caused by: TypeError: Failed to construct 'URL': Invalid URL
Better Auth 客户端内部用 new URL() 解析 baseURL,相对路径无法直接构造 URL。
解决方案:用 window.location.origin 拼接完整 URL:
export const authClient = createAuthClient({
baseURL: `${window.location.origin}/api/gate`,
})
本地开发展开为 http://localhost:5173/api/gate(Vite proxy 转发),线上展开为 https://your-domain.com/api/gate,无需环境变量。
最终生效的三个关键文件:
vercel.json
{
"buildCommand": "pnpm run check && pnpm exec esbuild server/src/app.ts --bundle --platform=node --format=cjs --outfile=dist/_server.cjs --external:@libsql/client && pnpm -F client run build",
"outputDirectory": "client/dist",
"rewrites": [
{ "source": "/api/(.*)", "destination": "/api/index" },
{ "source": "/((?!api/).*)", "destination": "/index.html" }
]
}
api/index.ts
import app from '../dist/_server.cjs';
export default app;
client/src/lib/auth-client.ts
import { createAuthClient } from 'better-auth/react';
export const authClient = createAuthClient({
baseURL: `${window.location.origin}/api/gate`,
});
Vercel 的 catch-all pattern 只能匹配单级路径(/api/foo),无法匹配多级(/api/foo/bar)。需要改用 api/index.ts + vercel.json rewrite 规则显式转发所有 /api/* 请求。
用 esbuild 打包为 CJS 格式(--format=cjs),原生依赖如 @libsql/client 用 --external 排除,并在根 package.json 声明为 dependency 让 Vercel 安装。
createAuthClient 的 baseURL 不支持相对路径,必须传完整 URL。用 window.location.origin 拼接可同时兼容本地开发和线上环境。
遇到类似的 Vercel 部署问题?联系我聊聊你的技术栈,看看能不能帮你少踩几个坑。
联系合作在为客户开发 Chrome 扩展数据采集工具时遇到此问题,记录根因与解法。
window.postMessage(data, '*') 会把消息广播给页面中所有 frame,包括第三方 iframe。如果你的 Chrome 扩展通过 postMessage 传递用户数据(如 memberId、业务报表),任何嵌入页面的 iframe 都能监听到。把 '*' 改为 window.location.origin 即可精确限定接收方。
Chrome 扩展的采集脚本通过 postMessage 将电商数据传递给 content script:
// ❌ 不安全:消息广播到所有 frame
window.postMessage({
type: 'CCL_SHOP_REPORT_DAILY',
memberId: 'b2b-2214126315258ad300', // 用户 ID
rows: [{ uv: 403, payAmt: 19478.47 }] // 业务数据
});
// 等同于 window.postMessage(data, '*')
页面上如果嵌入了第三方 iframe(广告、统计、社交插件),这些 iframe 的 message 事件监听器同样能收到这条消息。
postMessage 的第二个参数 targetOrigin 决定消息的接收范围:
| targetOrigin | 行为 |
|---|---|
'*' 或省略 | 广播到所有 frame,不检查来源 |
'https://example.com' | 只发送到 origin 为 https://example.com 的 frame |
window.location.origin | 只发送到当前页面同源的 frame |
省略第二个参数时,浏览器默认使用 '*'。这在 Chrome 扩展场景下尤其危险——扩展注入的脚本运行在电商平台页面,页面上可能有多个第三方 iframe。
// safePostMessage:强制使用 window.location.origin
function safePostMessage(data) {
window.postMessage(data, window.location.origin);
}
// 使用
safePostMessage({
type: 'CCL_SHOP_REPORT_DAILY',
subType: 'daily',
memberId: memberId,
rows: [row]
});
// content script 中监听消息
window.addEventListener('message', (event) => {
// ✅ 验证来源 origin
if (event.origin !== window.location.origin) return;
// ✅ 验证消息结构
if (!event.data || typeof event.data.type !== 'string') return;
switch (event.data.type) {
case 'CCL_SHOP_REPORT_DAILY':
handleDailyReport(event.data);
break;
case 'CCL_ITEM_WEEKLY_REPORT':
handleWeeklyReport(event.data);
break;
}
});
'*'只有一种场景安全:消息内容完全不含敏感信息,且接收方 origin 不可预知。例如纯 UI 状态通知("面板已打开")。即使如此,用 window.location.origin 也更安全。
postMessage 是它们通信的标准方式——务必保护好这条通道(遇到热重载后消息重复处理时需要手动清理旧监听器)event.origin 验证和发送端 targetOrigin 限定缺一不可,单向防护不完整chrome.runtime.sendMessage(参考 Service Worker Token 同步)而不是 postMessage在为客户构建电商数据分析平台时遇到此问题,记录根因与解法。
Drizzle ORM 的 sql 模板标签中,sql.join(values.map(v => sql(v))) 会把所有值参数化传递。如果 values 数组里混入了 SQL 表达式(如 date_trunc('week', '2026-05-17'::date)::date),PostgreSQL 会把它当成普通字符串解析,报 invalid input syntax for type date 错误。SQL 表达式必须用 sql.raw() 或单独写在模板外部。
电商数据采集流程:Chrome 扩展采集 → CCLHub 转发 → Analytics 写库。现象:
uv: 403, payAmt: 19478.47)uv: 0, pay_amt: 0.00-- 数据库实际数据
report_date | uv | pay_amt | reveal_cnt
-------------+-----+----------+------------
2026-05-12 | 392 | 7333.67 | 11879 -- 旧数据正常
2026-05-13 | 0 | 0.00 | 0 -- 新数据全零!
同时 Analytics 错误日志有:
PostgresError: invalid input syntax for type date:
"date_trunc('week', '2026-05-17'::date)::date"
原始代码混用了参数化值和 SQL 表达式:
// ❌ 问题代码
const insertVals: (string | number | null)[] = [
String(shop_id),
String(platform_id),
reportDate,
tenant_id,
`date_trunc('week', '${reportDate}'::date)::date`, // ← SQL 表达式
];
// sql.join 会把所有值参数化,包括 date_trunc 表达式
await db.execute(sql`
INSERT INTO table (..., week_start_date)
VALUES (${sql.join(insertVals.map(v => sql`${v}`), sql`,`)})
...
`);
生成的 SQL:
-- PostgreSQL 收到的 $5 参数值是字面字符串
INSERT INTO table (..., week_start_date)
VALUES ($1, $2, $3, $4, $5, ...)
-- $5 = "date_trunc('week', '2026-05-17'::date)::date" ← 被当字符串!
PostgreSQL 尝试把 "date_trunc('week', '2026-05-17'::date)::date" 解析为 date 类型 → 报错。
为什么数据是 0 而不是报错? 因为同一张表有独立的询盘写入(PARTIAL UPSERT),询盘 INSERT 成功创建了行(看板列默认值 0),日报 UPSERT 失败但没有回滚已存在的行。
把 SQL 表达式从参数化数组中分离出来,用 sql.raw() 或直接写在模板中:
// ✅ 修复:参数化值和 SQL 表达式分开
const insertCols = ['shop_id', 'platform_id', 'report_date', 'tenant_id'];
const insertVals: (string | number | null)[] = [
String(shop_id), String(platform_id), reportDate, tenant_id,
];
// 19 个数据列正常参数化
for (const [apiKey, dbCol] of Object.entries(DAILY_COLUMNS)) {
insertCols.push(dbCol);
insertVals.push(row[apiKey] != null ? String(row[apiKey]) : '0');
}
// week_start_date 用 SQL 表达式,不进参数化数组
await db.execute(sql`
INSERT INTO table (${sql.raw(insertCols.join(', '))}, week_start_date)
VALUES (
${sql.join(insertVals.map(v => sql`${v}`), sql`,`)},
date_trunc('week', ${reportDate}::date)::date -- ← 直接写在模板里
)
...
`);
关键区别:
| 写法 | Drizzle 处理方式 | PostgreSQL 收到 |
|---|---|---|
sql 模板插值 | 参数化($N) | 字符串字面量 |
sql.raw(expression) | 原样拼入 SQL | SQL 表达式 |
直接写在 sql 模板中 | 作为模板的一部分 | SQL 表达式 |
sql.raw() 存在 SQL 注入风险,不要用于用户输入。本例中 reportDate 来自内部 API,格式可控sql 模板标签会自动参数化所有插值——这是安全特性,但 SQL 函数调用不该被参数化在为客户构建 SaaS 认证系统时遇到此问题,记录根因与解法。
Node.js ES Module 中,import 语句在 dotenv.config() 之前执行。如果模块级代码读取 process.env.JWT_SECRET,拿到的是 undefined,导致 JWT 签名用 "undefined" 字符串作为密钥——不报错,但所有 token 校验都失败。解决方案:延迟初始化(lazy init)。
JWT 登录接口返回 200,但后续请求全部 401。排查发现:
jwtVerify() 验证process.env.JWT_SECRET,结果是 undefined// jwt.ts — 模块级代码
import crypto from 'crypto';
// ❌ 这行在 dotenv.config() 之前执行,JWT_SECRET 是 undefined
const SECRET = crypto.createSecretKey(
new TextEncoder().encode(process.env.JWT_SECRET)
);
最坑的是:不报错。new TextEncoder().encode(undefined) 会把字符串 "undefined" 编码成字节,生成一个合法但错误的密钥。
ES Module 的 import 是静态提升的:
// server.ts(入口文件)
import { router } from './routes/auth'; // ← 先执行
import { authenticateToken } from './middleware/auth'; // ← 先执行
dotenv.config(); // ← 后执行,但 import 链已经跑完了
执行顺序:
import,构建依赖图jwt.ts 的 const SECRET = ... 在这里执行)server.ts,执行 dotenv.config().env 才加载到 process.env所以 jwt.ts 模块级代码读到的 process.env.JWT_SECRET 是 undefined。
把密钥初始化从模块级移到函数内部,首次调用时才读取环境变量:
import crypto from 'crypto';
let _secret: crypto.KeyObject | null = null;
function getSecret(): crypto.KeyObject {
if (!_secret) {
const secretValue = process.env.JWT_SECRET;
if (!secretValue) {
throw new Error('JWT_SECRET 环境变量未设置');
}
_secret = crypto.createSecretKey(
new TextEncoder().encode(secretValue)
);
}
return _secret;
}
// 所有需要密钥的地方改用 getSecret()
export async function generateToken(payload: any): Promise<string> {
return new SignJWT(payload)
.setProtectedHeader({ alg: 'HS256' })
.sign(getSecret()); // ← 延迟到运行时读取
}
优势:不依赖入口文件的 import 顺序,任何调用时机都安全。
// server.ts — 确保这两行在所有 import 之前
import 'dotenv/config'; // 或 require('dotenv').config()
import express from 'express';
// ...其他 import
局限:如果有其他入口文件(如 cron job、worker)忘记加这行,问题复现。
require('dotenv').config() 只在 CommonJS 中能保证顺序;ES Module 中 import 始终先于运行时代码在为客户构建电商数据采集系统时遇到此问题,记录根因与解法。
1688 平台 Cookie 中 last_mid 和 unb 都包含用户 ID,但格式不同(b2b-xxx vs 纯数字)。数据库存的是 b2b- 前缀格式。原代码遍历 Cookie 数组时先匹配到了 unb,导致严格等于比较失败,所有采集数据写入时关联不到店铺,结果全是零值。
解法:把 Cookie key 优先级列表放在外层循环,Cookie 数组放在内层循环,确保高优先级的 key 先被查找。
1688 生意参谋日报采集后,数据库里看板数据全是 0,但询盘数据有值:
shop_id | report_date | reveal_cnt | uv | pay_amt | effective_inq_users
2 | 2026-05-13 | 0 | 0 | 0.00 | 56
2 | 2026-05-14 | 0 | 0 | 0.00 | 41
后端日志报错:No shop found for memberId: 2214126315258
但数据库里存的 platform_account_id 是 b2b-2214126315258ad300。
unb=2214126315258 # 纯数字
last_mid=b2b-2214126315258ad300 # b2b- 前缀 + 后缀
数据库映射表 shops.platform_account_id 存的是 b2b- 前缀格式。代码做的是严格等于匹配:
const match = mapping.find(m => m.platform_account_id === memberId);
// "2214126315258" !== "b2b-2214126315258ad300" → 匹配失败
原代码的外层循环是 Cookie 数组、内层是 key 列表:
// ❌ 错误:Cookie 数组在外层,key 优先级无效
var keys = ['last_mid', '__last_memberid__', 'unb'];
for (var i = 0; i < cookies.length; i++) { // 外层:Cookie
var pair = cookies[i].trim();
for (var k = 0; k < keys.length; k++) { // 内层:key
if (pair.indexOf(keys[k] + '=') === 0) {
return pair.substring(keys[k].length + 1);
}
}
}
document.cookie 的返回顺序不是固定的。如果 unb 的 Cookie 在数组中排在 last_mid 前面,就会先匹配到 unb,返回纯数字格式的 ID——last_mid 的优先级形同虚设。
交换循环层级:key 优先级列表放在外层,Cookie 数组放在内层:
// ✅ 正确:key 优先级列表在外层
var keys = ['last_mid', '__last_memberid__', 'unb'];
for (var k = 0; k < keys.length; k++) { // 外层:按优先级遍历 key
for (var i = 0; i < cookies.length; i++) { // 内层:在所有 Cookie 中查找
var pair = cookies[i].trim();
if (pair.indexOf(keys[k] + '=') === 0) {
return pair.substring(keys[k].length + 1);
}
}
}
key 列表按优先级在外层遍历,无论 document.cookie 返回顺序如何,始终先查找 last_mid,保证拿到和数据库格式一致的用户 ID,映射不会失败。
// 在浏览器控制台确认两个 Cookie 都存在
document.cookie.split(';')
.filter(c => /last_mid|unb/.test(c.trim()))
.map(c => c.trim())
// ['unb=2214126315258', 'last_mid=b2b-2214126315258ad300']
Chrome 扩展中使用 chrome.cookies.get() 读取 Cookie 时不存在此问题——它是按名称精确查询,天然支持优先级。但 document.cookie 字符串解析时务必注意循环层级。另外,如果你的扩展通过 postMessage 传递数据,也要注意 postMessage targetOrigin 的安全风险;如果热重载后消息被处理两次,需要手动管理监听器生命周期。
在为客户构建 SaaS 数据分析平台时遇到此问题,记录根因与解法。
Node.js ESM 模式下,import('./path/to/module') 不会自动解析 ./path/to/module.js。TypeScript 编译的 dist 产物如果遗漏 .js 后缀,模块加载时抛出 ERR_MODULE_NOT_FOUND。如果这个 import 在延迟逻辑中(如定时器、条件分支),应用启动正常但运行一段时间后崩溃,PM2 表现为重启计数飙升。
解法:确保所有 ESM 动态 import 路径包含 .js 扩展名,并在 build 流程中自动修复。
PM2 状态显示应用不断重启:
│ name │ ↺ │ status │ uptime │
│ analytics-api │ 9 │ online │ 28m │
错误日志每几分钟重复:
Error [ERR_MODULE_NOT_FOUND]: Cannot find module '/app/dist/domains/video/cleanup'
imported from /app/dist/server.js
但文件实际存在:
$ ls dist/domains/video/
cleanup.js executor.js queue.js
Node.js 的 CommonJS (require()) 会自动尝试 .js、.json 等后缀。ESM (import) 不会。
// ❌ ESM 模式下找不到模块
import('./domains/video/cleanup')
// Node.js 查找: ./domains/video/cleanup (精确路径,无扩展名)
// 实际文件: ./domains/video/cleanup.js
// ✅ 必须带 .js 后缀
import('./domains/video/cleanup.js')
这个 import 在定时器中延迟执行:
// server.ts — 启动时不立即执行
import('./domains/video/cleanup.js').then(({ startCleanupScheduler }) => {
startCleanupScheduler(); // 几秒后才触发
});
应用启动成功(DB 连接、端口监听都正常),定时器触发时 import 报错 → 进程崩溃 → PM2 重启 → 再次启动成功 → 定时器再次触发 → 再次崩溃。形成崩溃循环。
旧版本的 dist 产物是通过手动 build 生成的,build 脚本包含 .js 后缀修复步骤。某次部署时跳过了 build 后的修复步骤,直接用 tsc 输出的产物部署,tsc 不修改 import 路径。
.js 后缀TypeScript 官方推荐:即使在 .ts 文件中,也写 .js 后缀:
// ✅ TypeScript 源码中也写 .js
import('./domains/video/cleanup.js').then(({ startCleanupScheduler }) => {
startCleanupScheduler();
});
在 build 流程中添加后缀修复脚本,自动为没有扩展名的 import 补上 .js:
{
"scripts": {
"build": "tsc && node fix-imports.js"
}
}
fix-imports.js 核心逻辑:
import { readFileSync, writeFileSync, readdirSync } from 'fs';
import { join } from 'path';
function fixImports(dir) {
for (const file of readdirSync(dir, { withFileTypes: true })) {
const fullPath = join(dir, file.name);
if (file.isDirectory()) {
fixImports(fullPath);
} else if (file.name.endsWith('.js')) {
let content = readFileSync(fullPath, 'utf8');
// 修复动态 import: from 'xxx' → from 'xxx.js'
const fixed = content.replace(
/import\(['"](\.[^'"]+)['"]\)/g,
(match, path) => path.endsWith('.js') ? match : match.replace(path, path + '.js')
);
// 修复静态 import: from 'xxx' → from 'xxx.js'
const fixed2 = fixed.replace(
/from\s+['"](\.[^'"]+)['"]/g,
(match, path) => path.endsWith('.js') ? match : match.replace(path, path + '.js')
);
if (fixed2 !== content) {
writeFileSync(fullPath, fixed2);
}
}
}
}
build 流程自动为所有相对路径 import 补全 .js 后缀,TypeScript 源码保持无后缀写法,ESM 部署不再出现模块找不到的错误。
"type": "module" 或 .mjs 文件)下出现,CommonJS 不受影响import ... from './foo')同样受此限制,不仅限于动态 import();ESM 的 import 提升还会导致另一个常见问题——dotenv 在 import 链之后才执行,环境变量读不到tsx、ts-node 开发时不报错(它们会自动解析后缀),但 node dist/server.js 生产运行时出错在为客户构建电商数据采集 Chrome 扩展时遇到此问题,记录根因与解法。
WXT 框架 HMR 热重载时,content script 重新执行但旧的 window.addEventListener('message', ...) 不会被清除。每热重载一次就多一个监听器实例,导致每条 postMessage 被处理 N 次。
解法:注册新监听器前,从 window 变量取出旧监听器引用并 removeEventListener。
浏览器控制台显示同一条消息被两个不同实例捕获:
content.js:114 [CCL] CCL_SHOP_REPORT_DAILY daily caught - 实例: nxctn6
content.js:2 [CCL] CCL_SHOP_REPORT_DAILY daily caught - 实例: t6jce7
每条 postMessage 被处理两次,导致后台发送重复请求。
WXT (基于 Vite 的 Chrome 扩展框架) 开发模式下,修改 content script 后触发 HMR:
window.addEventListener('message', messageListener) 被注册结果:window 上挂载了多个独立的 message 监听器,每条 postMessage 触发所有实例。
在 content script 入口处,注册新监听器前移除旧的:
const instanceId = Math.random().toString(36).slice(2, 8);
// 取出旧监听器引用
const prevListener = (window as any).__cclMessageListener;
if (prevListener) {
window.removeEventListener('message', prevListener);
}
// 定义新监听器
const messageListener = (event: MessageEvent) => {
// ... 处理逻辑
};
// 存储当前引用(供下次 HMR 取用)
(window as any).__cclMessageListener = messageListener;
// 注册
window.addEventListener('message', messageListener);
关键点:removeEventListener 必须传入和 addEventListener 相同的函数引用。把函数存到 window 变量上,下次 HMR 时就能取出旧引用并正确移除。这样无论热重载多少次,始终只有一个活跃的 message 监听器。
如果你同时遇到 postMessage 的 targetOrigin 安全问题或Cookie 取值导致数据全零,建议一并排查。
window 上的变量在页面刷新前一直存在,HMR 只替换脚本模块不清除 window 属性