1. The Node.js Platform
本章建立后续所有章节的基础:理解 Node.js 为什么这样工作,以及为什么 Node.js 的设计方式与传统服务器端编程不同。原书将这一章的核心归结为 Node way、Reactor Pattern、V8 + libuv + Node.js Core。
1.1 The Node.js philosophy
Small core
Node.js Core 应尽可能保持小,把大量能力放到 userland。
核心思想:
- Core 只提供稳定、基础、通用的能力。
- 更具体的解决方案交给 npm/userland。
- 避免核心过于庞大导致演进缓慢。
- 让社区可以快速尝试不同方案。
思想本质:
Core 提供基础设施,生态系统提供具体解决方案。
这也是 Node.js 生态高度繁荣的重要原因。
Small modules
模块应该:
- 小;
- 单一职责;
- 有明确边界;
- 易理解;
- 易测试;
- 易复用。
其思想来自 Unix:
- Small is beautiful。
- Make each program do one thing well。
Node.js 的 npm 生态允许大量小模块共存,因此”小模块”在 Node.js 中比传统平台更加可行。
Small surface area
模块不仅应该小,而且应该拥有尽可能小的公开接口。
核心原则:
- 不暴露内部实现。
- 对外只提供真正需要的功能。
- 优先暴露 function,而不是为了扩展性而暴露复杂 class hierarchy。
- “可使用”通常比”可继承、可扩展”更重要。
结果:
API 越小:
使用方式越明确 → 错误使用越少 → 实现越简单 → 维护越容易。
Simplicity and pragmatism
Node.js 倾向于:
- KISS;
- “worse is better”;
- 尽快得到一个足够好的解决方案;
- 避免为了理论上的完美而引入巨大复杂度。
因此:
简单、可维护、实际可用,通常比理论上完美更加重要。
这也是为什么传统 GoF Pattern 在 Node.js 中经常可以被极大简化。
1.2 How Node.js works
I/O is slow
相对于 CPU 和 RAM,I/O 延迟非常高。
典型层次:
CPU/RAM → 文件系统 → 网络 → 人类输入
Node.js 的整个异步架构,本质上是在解决:
如何让 CPU 不因为等待 I/O 而闲置。
Blocking I/O
阻塞 I/O:
调用 I/O
↓
线程等待
↓
I/O 完成
↓
继续执行
传统解决方法:
一个连接 → 一个线程
问题:
- 线程占用内存;
- 上下文切换有成本;
- 大部分时间线程实际上在等待 I/O;
- 高并发时资源浪费严重。
Non-blocking I/O
非阻塞 I/O:
- I/O 调用立即返回;
- 如果暂时没有数据,则返回”暂不可用”;
- 应用可以继续处理其他工作。
简单的 busy-wait:
循环检查资源
↓
没有数据 → 再检查
↓
有数据 → 处理
但这种方式浪费 CPU。
Event demultiplexing
更高效的方法:
Application
↓
Event Demultiplexer
↓
等待多个资源
↓
某个资源就绪
↓
返回事件
↓
执行对应 handler
它能够:
- 在一个线程中管理大量 I/O;
- 不需要轮询所有资源;
- 只有真正有事件时才工作。
The reactor pattern
Reactor Pattern 是 Node.js 异步体系的核心。
基本模型:
Resource
↓
注册事件
↓
Reactor / Event Demultiplexer
↓
Event Loop
↓
Handler
核心流程:
- 注册需要监听的 I/O。
- Reactor 等待事件。
- I/O 就绪。
- Reactor 获取事件。
- 调用对应 handler。
- 继续等待下一批事件。
所以 Node.js 的”单线程”并不等于”一次只能处理一个请求”。
它真正的模型是:
JavaScript 执行本身是单线程的,但 I/O 并发由操作系统、libuv 和事件循环共同完成。
Libuv, the I/O engine of Node.js
Node.js 的异步能力并不是 JavaScript 自己实现的,而是建立在 libuv 上。
libuv 负责:
- 事件循环;
- 操作系统异步 I/O;
- 文件系统;
- 网络;
- 定时器;
- 某些无法直接异步执行的操作的线程池。
1.3 The recipe for Node.js
Node.js 可以理解为几个核心部件组合而成:
V8
+
libuv
+
Node.js Core APIs
+
Bindings
=
Node.js
V8
负责:
- JavaScript 执行;
- JIT;
- Garbage Collection;
- 内存管理。
libuv
负责底层异步 I/O 与 event loop。
Bindings
把底层能力暴露给 JavaScript。
Node.js Core Library
提供:
- fs;
- http;
- stream;
- crypto;
- events;
- child_process;
- 等高级 API。
1.4 JavaScript in Node.js
Run the latest JavaScript with confidence
浏览器需要面对:
Chrome
Firefox
Safari
不同版本
不同能力
Node.js 服务端通常可以控制运行环境,因此可以:
- 指定 Node.js 版本;
- 使用现代 JavaScript;
- 减少兼容性代码;
- 减少 transpiler/polyfill 依赖。
The module system
Node.js 没有浏览器的:
- DOM;
- window;
- document。
但拥有:
- filesystem;
- network;
- process;
- OS;
- native modules。
Full access to operating system services
Node.js 的能力远大于浏览器,因为它是服务器端运行环境。
也正因为如此:
Node.js 应用的安全边界比浏览器 JavaScript 更重要。
Running native code
Node.js 可以通过:
- Native Addons;
- N-API;
- WebAssembly;
调用/运行非 JavaScript 代码。
本章核心结论
Node.js = 小核心 + 小模块 + 小 API + 简单务实 + Reactor/Event Loop + V8 + libuv。
作者的总结也明确把这些作为本章核心。
2. The Module System
本章完整讨论 CommonJS 与 ESM,以及模块解析、缓存、循环依赖和互操作。原书把这两种 module system 作为 Node.js 中的两套核心模块机制。
2.1 The need for modules
模块解决的问题:
- namespace;
- 封装;
- 依赖管理;
- 复用;
- 可测试性;
- 大型系统拆分。
没有 module system 时:
所有代码共享 global scope
↓
命名冲突
↓
隐式依赖
↓
难以维护
所有代码共享 global scope
↓
命名冲突
↓
隐式依赖
↓
难以维护2.2 Module systems in JavaScript and Node.js
JavaScript 主要经历:
Global scripts
↓
IIFE(Immediately Invoked Function Expression)
↓
Revealing Module Pattern
↓
CommonJS / AMD(Asynchronous Module Definition)
↓
ES Modules
Node.js 历史上主要使用 CommonJS,现代 Node.js 同时支持 ESM。
2.3 The module system and its patterns
The revealing module pattern
const myModule = (() => {
const privateFoo = () => {}
const privateBar = []
const exported = {
publicFoo: () => {},
publicBar: () => {}
}
return exported
})() // once the parenthesis here are parsed, the function will be invoked
console.log(myModule)
console.log(myModule.privateFoo, myModule.privateBar)
通过 closure 隐藏 private state:
private data
↓
closure
↓
只暴露需要的 API
核心思想:
内部状态私有,对外只暴露有限接口。
2.4 CommonJS modules
A homemade module loader
function loadModule(filename, module, require) {
const wrappedSrc =
`(function (module, exports, require) {
${fs.readFileSync(filename, 'utf8')}
})(module, module.exports, require)`
eval(wrappedSrc)
}
function require(moduleName) {
console.log(`Require invoked for module: ${moduleName}`)
const id = require.resolve(moduleName)
if (require.cached[id]) {
return require.cached[id].exports
}
// module metadata
const module = {
exports: {},
id
}
// update the cache
require.cache[id] = module
// load the module
loadModule(id, module, require)
// return exported variables
}
require.cache = {}
require.resolve = (moduleName) => {
// resolve a full module id from the moduleName
}
理解 CommonJS 的关键是理解 require() 可以看作:
resolve
↓
load
↓
wrap
↓
execute
↓
cache
↓
return exports
Node.js 实际上会对模块代码进行类似 wrapper 的封装,使:
- module;
- exports;
- require;
- __filename;
- __dirname
成为模块局部变量。
Defining a module
CommonJS:
module.exports = ...
真正导出的对象是:
module.exports
module.exports versus exports
这是最容易出错的知识点之一。
初始关系类似:
exports === module.exports
但:
exports.foo = ...
是在修改原对象。
而:
exports = ...
只是改变局部变量引用。
所以:
需要整体替换导出对象时使用
module.exports。
The require function is synchronous
CommonJS:
const foo = require('./foo')
是同步加载。
因此:
- 模块解析发生在当前执行流程;
- 加载/执行模块会阻塞当前 JavaScript 执行;
- CommonJS 天然适合启动阶段加载依赖。
The resolving algorithm
require() 大致要判断:
- Core module
- File module
- Package module
这解释了为什么:
require('fs')
require('./foo')
require('some-package')
会走不同的解析路径。
myApp
├── foo.js
└── node_modules
├── depA
│ └── index.js
├── depB
│ ├── bar.js
│ └── node_modules
│ └── depA
│ └── index.js
└── depC
├── foobar.js
└── node_modules
└── depA
└── index.js
- Calling
require('depA')from/myApp/foo.jswill load/myApp/node_modules/depA/index.js - Calling
require('depA')from/myApp/node_modules/depB/bar.jswill load/myApp/node_modules/depB/node_modules/depA/index.js - Calling
require('depA')from/myApp/node_modules/depC/foobar.jswill load/myApp/node_modules/depC/node_modules/depA/index.js
The module cache
模块第一次被加载:
resolve
→ execute
→ cache
之后再次:
require()
→ cache
→ 返回同一模块实例
因此:
CommonJS module cache 天然形成”进程级共享实例”。
这也是 Node.js Singleton 实现非常简单的重要原因。
但必须注意:
Singleton / module cache 是 process-local,并不意味着整个分布式系统只有一个实例。
Circular dependencies
循环依赖:
A → B
↑ ↓
└───┘
CommonJS 的一个关键问题:
- 模块执行具有顺序;
- 当模块尚未完成初始化时,另一个模块可能已经拿到它的”部分 exports”。
因此可能看到:
partial exports
这也是 CommonJS circular dependency 容易出现 undefined/部分对象的原因。
2.5 Module definition patterns
Named exports
导出多个明确能力。
适合:
- utility;
- 多个相关函数;
- public API。
Exporting a function
非常符合 Node.js philosophy:
一个模块可以只做一件事情。
Exporting a class
适合:
- 有明确对象生命周期;
- 需要实例化;
- 维护对象状态。
但 Node.js 并不鼓励为了 OOP 而强行使用 class。
Exporting an instance
直接导出实例:
module
↓
instance
常用于:
- shared state;
- connection;
- singleton-like component。
Modifying other modules or the global scope
包括:
- monkey patch;
- 修改 global;
- 修改第三方模块。
风险很高:
- hidden dependency;
- 难以测试;
- 难以推导行为;
- 模块之间产生隐式耦合。
所以:
除非有非常明确的原因,否则避免修改其他模块或 global scope。
2.6 ESM: ECMAScript modules
Using ESM in Node.js
ESM 采用:
import ...
export ...
与 CommonJS 有明显不同。
Named exports and imports
支持:
export const foo = ...
export function bar () {}
对应:
import { foo, bar } from './module.js'
核心优点:
- 静态结构;
- dependency graph 可提前分析;
- 更适合工具链;
- tree shaking 更容易。
Default exports and imports
支持一个主要默认导出:
export default ...
适合一个模块只有一个主要概念。
Mixed exports
可以同时:
- named exports;
- default export。
但从 API 设计角度应避免让模块接口过于复杂。
Module identifiers
ESM 的模块标识规则与 CommonJS 不同。
重要区别:
ESM import 通常需要明确文件扩展名,而 CommonJS
require()在很多场景下可以省略。
Async imports
import() 是动态导入:
import()
↓
Promise
↓
模块加载完成
适合:
- lazy loading;
- 按需加载;
- runtime 决定模块;
- 减少初始加载成本。
2.7 Module loading in depth
Loading phases
ESM 的关键机制:
Phase 1 — Parsing / Construction
以深度优先的方式发现所有 import。
建立:
dependency graph
Phase 2 — Instantiation
创建 import/export 的绑定关系。
此时:
- 建立引用;
- 还没有真正执行模块代码。
Phase 3 — Evaluation
执行模块代码,给绑定赋实际值。
所以:
Parsing
→ Instantiation
→ Evaluation
这个模型是理解 ESM circular dependency 的关键。
Read-only live bindings
ESM 的 import 是:
read-only live binding
即:
- consumer 不能重新给 imported binding 赋值;
- 但 exporter 模块内部修改变量时,consumer 可以看到更新。
这与 CommonJS 很不同。
CommonJS 更接近:
exports object 的浅复制/对象引用语义。
ESM 更接近:
真正的符号绑定关系。
Circular dependency resolution
ESM 能比 CommonJS 更系统地处理 circular dependency,因为:
- 解析阶段先构建完整 dependency graph;
- instantiation 阶段建立所有 bindings;
- evaluation 阶段再执行代码。
因此循环依赖下可以保留更加完整的引用关系。
Modifying other modules
import fs, { readFileSync } from 'fs'
import { syncBuiltinESMExports } from 'module'
fs.readFileSync = () => Buffer.from('Hello, ESM')
syncBuiltinESMExports()
console.log(fs.readFileSync === readFileSync) // true
ESM 更严格地强调:
- import binding 是 read-only;
- module interface 应当由模块自己定义;
- 不应该依赖 monkey patch 式修改其他模块。
2.8 ESM and CommonJS differences and interoperability
ESM runs in strict mode
ESM 自动 strict mode。
Missing references in ESM
ESM 没有 CommonJS 的:
require
exports
module.exports
__filename
__dirname
解决方法
import { fileURLToPath } from 'url'
import { dirname } from 'path'
const __filename = fileURLToPath(import.meta.url)
const __dirname = dirname(__filename)
import { createRequire } from 'module'
const require = createRequire(import.meta.url)
this 在 ES 中是未定义的
// this.js - ESM
console.log(this) // undefined
// this.cjs - CommonJS
console.log(this === exports) // true
Interoperability
现代 Node.js 可以让 CommonJS 与 ESM 互操作,但两种模块模型本质上不同。
ESM 中导入 CommonJS(仅限默认导出)
import packageMain from 'commonjs-package' // Works
import { method } from 'commonjs-package' // Errors
ESM 中导入 JSON 数据
import { createRequire } from 'module'
const require = createRequire(import.meta.url)
const data = require('./data.json')
console.log(data)
因此设计时要清楚:
“能互相调用”不等于”两个 module system 语义完全相同”。
本章核心结论
模块系统的核心不是记住 require/import 语法,而是理解:
Module
├─ dependency
├─ encapsulation
├─ initialization
├─ caching
├─ resolution
├─ execution order
└─ interoperability
作者最终要求掌握 CommonJS 和 ESM 两套体系,而后续书中主要采用 ESM。
3. Callbacks and Events
这一章建立 Node.js 异步编程的两个基本原语:
Callback
EventEmitter
原书称它们为 Node.js asynchronous infrastructure 的两个支柱。
3.1 The Callback pattern
Continuation-Passing Style
CPS 的思想:
函数不直接返回最终结果,而是把”接下来要执行什么”作为 callback 传进去。
普通:
result = operation()
CPS:
operation(input, callback)
operation(input, callback)Synchronous CPS
callback 立即调用:
operation()
↓
callback()
仍然是同步执行。
Asynchronous CPS
callback 在未来执行:
operation()
↓
return
↓
event loop
↓
callback()
Node.js 绝大多数 I/O API 使用这种方式。
Non-CPS callbacks
callback 不一定代表异步 continuation。
因此:
看到 callback 并不能自动推断函数是异步的。
3.2 Synchronous or asynchronous?
这是 Node.js callback API 最危险的问题之一。
一个函数可能:
某些情况同步
某些情况异步
例如:
cache hit → synchronous
cache miss → asynchronous
这种 API 会产生极难推理的问题。
An unpredictable function
调用方无法确定 callback 什么时候运行。
导致:
- execution order 不明确;
- stack behavior 不一致;
- error handling 复杂;
- 测试不稳定。
Unleashing Zalgo
Zalgo 指:
一个 API 在某些情况下同步调用 callback,在另一些情况下异步调用 callback。
这是非常危险的 API 设计。
原则:
异步 API 应该保证 callback 始终异步执行。
Using synchronous APIs
有时同步 API 本身并不是问题。
适合:
- 启动阶段;
- CLI;
- 小工具;
- 明确不需要并发的代码。
但在服务器请求处理中使用大量同步 I/O,会阻塞 event loop。
Guaranteeing asynchronicity with deferred execution
可以通过:
process.nextTick()setImmediate()- 其他异步调度
把 callback 推迟执行。
但要知道:
process.nextTick()的优先级非常高,大量递归使用可能导致 I/O starvation。
3.3 Node.js callback conventions
Node.js callback 约定:
The callback comes last
doSomething(arg1, arg2, callback)
Any error always comes first
callback(err, result)
即:
err == null
↓
success
否则:
error
errorPropagating errors
异步 callback 中:
lower-level error
↓
callback(err)
↓
upper-level callback
每一层必须正确传递 error。
典型错误:
- 忽略 error;
- 忘记调用 callback;
- callback 两次;
- error 被吞掉。
Uncaught exceptions
异步 callback 中抛出的异常不会自动像同步调用那样自然向调用方传播。
因此:
asynchronous API 中必须有清晰的 error channel。
3.4 The Observer pattern
Observer:
Subject
↓
notify
↓
Observers
Node.js 对应:
EventEmitter
EventEmitterThe EventEmitter
核心 API:
on()
once()
emit()
removeListener()
on()
once()
emit()
removeListener()Creating and using EventEmitter
适合:
一个操作会产生 多个事件/多个通知。
例如:
start
progress
data
error
end
start
progress
data
error
endPropagating errors
EventEmitter 有一个特殊事件:
error
如果没有 listener,可能导致未捕获异常。
因此:
使用 EventEmitter 时必须认真设计 error event。
Making any object observable
可以:
- 继承 EventEmitter;
- 组合 EventEmitter;
- 让对象拥有事件通知能力。
重点不是继承本身,而是:
将”状态变化”与”消费者”解耦。
EventEmitter and memory leaks
listener 会持有引用。
如果:
listener 注册
↓
永远没有 remove
可能产生:
- memory leak;
- listener 数量增长;
- 不必要的 callback 执行。
once() 可以自动移除一次性 listener,但如果事件永远不发生,listener 仍可能一直存在。
Synchronous and asynchronous events
同步事件:
emit()
↓
listener immediately runs
异步事件:
emit/schedule
↓
event loop
↓
listener runs
同步事件的问题是:
如果 listener 在事件产生之后才注册,就会错过事件。
因此 EventEmitter 通常更适合异步事件流。
EventEmitter versus callbacks
最重要的判断标准:
callback 用于返回一个结果;event 用于通知”发生了某件事情”。
也就是:
一次结果 → callback
多次通知 → event
一次结果 → callback
多次通知 → eventCombining callbacks and events
两者可以结合:
EventEmitter → progress/events
Callback → final result
这是实际 Node.js API 中非常常见的设计。
4. Asynchronous Control Flow Patterns with Callbacks
重点从”如何异步执行一个操作”提升到:
如何组织多个异步操作。
原书特别强调 sequential、parallel、limited parallel 三类控制流。
4.1 The difficulties of asynchronous programming
异步程序最大的复杂度不是”异步”本身,而是:
多个异步任务
+
依赖关系
+
错误
+
并发
+
完成条件
多个异步任务
+
依赖关系
+
错误
+
并发
+
完成条件Creating a simple web spider
Spider 是本章贯穿案例,用来演示:
- recursion;
- sequential;
- parallel;
- race conditions;
- concurrency limit。
Callback hell
典型:
callback
└─ callback
└─ callback
└─ callback
问题:
- indentation;
- readability;
- error propagation;
- control flow 难理解。
但真正的问题不是”嵌套太多”这么简单。
本质是:
控制流与业务逻辑混杂。
4.2 Callback best practices and control flow patterns
Callback discipline
基本原则:
一个 callback 只调用一次
避免:
success → callback()
error → callback(err)
两条路径同时执行。
尽早 return
避免:
if (...) {
...
} else {
...
}
可以:
if (err) return callback(err)
减少 nesting。
明确所有异步出口
确保:
- success 有 callback;
- error 有 callback;
- empty case 有 callback;
- early return 也有 callback。
4.3 Sequential execution
适用于:
A → B → C
B 依赖 A,C 依赖 B。
Executing a known set of tasks in sequence
如果任务集合已知:
task1
↓
task2
↓
task3
核心是维护:
current index
每完成一个任务再启动下一个。
Sequential iteration
例如:
process(items)
需要保证:
item1 完成
↓
item2
↓
item3
这适合:
- 有顺序要求;
- 任务之间存在资源限制;
- 不希望瞬间产生大量并发。
4.4 Parallel execution
如果任务互不依赖:
┌─ Task A ─┐
Start ─┼─ Task B ─┼─> All done
└─ Task C ─┘
性能通常更好。
核心难点:
怎么知道所有任务都结束?
经典方法:
completed counter
+
final callback
completed counter
+
final callbackWeb spider version 3
将:
A → B → C
改成:
A ─┐
B ─┼→ wait all
C ─┘
可以显著减少总耗时。
4.5 The pattern / Fixing race conditions with concurrent tasks
并行执行会产生 race condition。
例如:
Task A 修改状态
Task B 读取状态
如果顺序不确定,就可能出现错误。
因此:
并发不是简单地”同时执行所有任务”,必须明确共享状态和完成条件。
4.6 Limited parallel execution
真正生产环境中通常不能:
100,000 tasks → 同时启动
因为会导致:
- connection exhaustion;
- memory pressure;
- CPU overload;
- downstream overload。
所以需要:
max concurrency = N
max concurrency = NLimiting concurrency
典型模型:
Queue
↓
N workers
↓
Task
保持:
activeTasks <= N
activeTasks <= NGlobally limiting concurrency
有时限制不只是某一次操作,而是:
整个应用对某个资源的全局并发量。
例如:
DB max 10
HTTP max 20
file operations max 5
需要共享 concurrency controller。
4.7 The async library
本章最后强调:
实际生产环境不要轻易自己实现所有异步控制流算法。
成熟库已经解决:
- series;
- parallel;
- waterfall;
- queue;
- retry;
- concurrency;
- error handling。
生产环境除非有特殊需求,否则应该优先使用成熟、经过验证的实现。
5. Asynchronous Control Flow Patterns with Promises and Async/Await
本章进入现代 JavaScript 异步核心:
Promise
+
async/await
原书强调它们能够显著简化 serial、parallel、error handling,同时注意 forEach()、return await 和递归 Promise chain 等陷阱。
5.1 Promises
What is a promise?
Promise 表示:
一个异步操作未来的结果。
三种状态:
Pending
↓
Fulfilled
Pending
↓
Rejected
一旦 settled:
fulfilled / rejected
就不能再次改变状态。
Promises/A+ and thenables
Promise/A+ 强调统一的:
then(onFulfilled, onRejected)
核心能力。
Thenable 指拥有:
then(...)
接口的对象,可以参与 Promise resolution。
5.2 The promise API
最重要:
then()
catch()
finally()
以及:
Promise.resolve()
Promise.reject()
Promise.all()
...
最关键的 Promise 特性
p2 = p1.then(...)
then() 会返回另一个 Promise。
这使得:
A
↓
then
↓
B
↓
then
↓
C
可以形成 Promise Chain。
5.3 Creating a promise
Promise executor:
new Promise((resolve, reject) => {
...
})
resolve
成功。
reject
失败。
重要:
executor 本身是立即执行的,而 Promise 的结果可以稍后 settle。
5.4 Promisification
把:
callback(err, result)
转换成:
Promise
例如:
old API
↓
promisify
↓
Promise API
好处:
- 更容易组合;
- 更容易统一 error handling;
- 更容易使用 async/await。
5.5 Sequential execution and iteration
Promise chain 非常适合:
A
↓
B
↓
C
例如:
doA()
.then(doB)
.then(doC)
这就是 Promise 的经典 sequential pattern。
5.6 Parallel execution
没有依赖时:
await Promise.all([
taskA(),
taskB(),
taskC()
])
即:
A ─┐
B ─┼→ Promise.all
C ─┘
优势:
- 简洁;
- 错误传播统一;
- 不需要自己管理 counter。
5.7 Limited parallel execution
Promise.all() 最大的问题:
它会一次启动所有任务。
因此需要:
TaskQueue
Producer
Consumer
Concurrency Limit
这是 Producer-Consumer pattern。
5.8 Implementing the TaskQueue class with promises
典型设计:
taskQueue
consumer
concurrency
消费者数量决定并发量:
concurrency = N
队列有任务:
consumer → task
队列为空:
consumer → await
这是 Node.js 非常重要的并发控制思想。
5.9 Async/await
Async functions and await
async function:
总是返回 Promise。
await:
等待 Promise settle,并把控制权交还 event loop。
代码因此接近同步形式:
A
↓
await
↓
B
↓
await
↓
C
但底层仍然是异步的。
5.10 Error handling with async/await
核心:
try {
await something()
} catch (err) {
...
}
相比 callback:
callback error forwarding
更加直观。
A unified try…catch experience
async/await 最大价值之一:
asynchronous error handling 可以重新使用熟悉的
try/catch模型。
The “return” versus “return await” trap
两者在很多情况下看似一样:
return promise
与:
return await promise
但在 try/catch/finally 等场景中行为可能不同。
核心认识:
return await会让当前 async function 等待 Promise settle 后再进行返回流程,因此会影响异常捕获边界。
不要机械地认为:
return await = always bad
要根据 error-handling 语义判断。
5.11 Sequential execution and iteration
最自然的方式:
for (const item of items) {
await process(item)
}
这样天然保证:
item1 → item2 → item3
item1 → item2 → item35.12 Antipattern — async/await + Array.forEach
这是全书一个非常重要的实际陷阱。
不要:
items.forEach(async item => {
await process(item)
})
因为:
forEach()不等待 callback 返回的 Promise。
所以外层不会等待所有任务完成。
需要 serial:
for...of + await
需要 parallel:
Promise.all(items.map(...))
Promise.all(items.map(...))5.13 Parallel execution
典型:
await Promise.all(items.map(process))
await Promise.all(items.map(process))5.14 Limited parallel execution
需要:
queue
+
consumer pool
例如:
max concurrency = 10
而不是:
Promise.all(100000 tasks)
Promise.all(100000 tasks)5.15 Infinite recursive promise resolution chains
一个非常容易忽略的问题:
Promise
↓
then
↓
Promise
↓
then
↓
Promise
...
大量微任务可能持续占据 event loop。
因此:
Promise 本身是异步控制工具,但不意味着无限 Promise recursion 是安全的。
6. Coding with Streams
这是 Node.js 最重要的章节之一。书中明确认为 Streams 是处理 binary data、strings、objects 的核心模式,并强调 stream composition。
6.1 Discovering the importance of streams
Buffering versus streaming
Buffering:
全部数据
↓
内存
↓
处理
Streaming:
chunk1 → process
chunk2 → process
chunk3 → process
chunk1 → process
chunk2 → process
chunk3 → processSpatial efficiency
Streaming 最大优势:
不需要一次把全部数据放入内存。
所以特别适合:
- 大文件;
- HTTP;
- 视频;
- 数据处理;
- 大量记录。
Time efficiency
Streaming 可以:
数据到达一点就开始处理一点。
因此:
开始处理时间
与:
整个数据读取完成时间
可以解耦。
Composability
Stream 最重要的特性之一:
Readable
↓
Transform
↓
Transform
↓
Writable
每个组件只负责一件事情。
这与 Node.js small modules philosophy 完美契合。
6.2 Getting started with streams
Anatomy of streams
主要四类:
Readable
Writable
Duplex
Transform
另有:
PassThrough
PassThrough6.3 Readable streams
Readable 表示:
数据来源。
例如:
- file;
- socket;
- stdin;
- HTTP request。
Reading from a stream
两种主要模式:
Flowing mode
数据自动通过 data 事件流动。
Non-flowing / paused mode
消费者主动读取。
Implementing Readable streams
核心方法:
_read()
通过 push 数据进入 Readable。
重要思想:
Readable 不应该一次生成所有数据,而应该按需产生。
6.4 Writable streams
Writable 表示:
数据目的地。
例如:
- file;
- socket;
- stdout;
- database。
核心:
write(chunk)
end()
write(chunk)
end()Backpressure
这是 Streams 最重要的知识点之一。
当:
Producer speed
>
Consumer speed
就会产生:
buffer accumulation
Writable write() 的返回值就是 backpressure 信号。
大致逻辑:
write() === false
↓
pause producer
↓
wait drain
↓
resume
所以:
Backpressure 是生产者与消费者速度不匹配时的流量控制机制。
Implementing Writable streams
核心方法:
_write(chunk, encoding, callback)
调用 callback 表示当前 chunk 已处理。
6.5 Duplex streams
同时:
Readable
+
Writable
典型:
TCP socket
TCP socket6.6 Transform streams
Transform:
input
↓
transform
↓
output
典型:
- gzip;
- encryption;
- parsing;
- filtering;
- aggregation。
Filtering and aggregating data
Transform 不一定是一进一出:
input chunks
↓
filter
↓
aggregate
↓
output
因此可以做:
- filtering;
- grouping;
- statistics;
- parsing。
6.7 PassThrough streams
PassThrough:
数据基本原样通过。
常用于:
- monitoring;
- instrumentation;
- debugging;
- branching;
- observing stream。
6.8 Observability
可以通过 PassThrough 或其他机制:
Readable
↓
Observable layer
↓
Writable
统计:
- bytes;
- chunks;
- throughput;
- timing。
6.9 Late piping
如果太晚建立 pipe,可能:
数据已经开始流动,消费者可能错过数据。
因此要理解:
- flowing mode;
- listener 注册时间;
- pipe 注册时间。
6.10 Lazy streams
Stream 与 generator/iterator 配合可以实现真正 lazy 的数据来源:
不是:
Array → Stream
而是:
Generator → Stream
这样不会提前产生全部数据。
6.11 Connecting streams using pipes
基本:
readable.pipe(transform).pipe(writable)
readable.pipe(transform).pipe(writable)Pipes and error handling
一个常见问题:
pipe 链中的错误并不会自动等于”所有相关资源都被安全处理”。
因此需要认真处理:
- error;
- close;
- cleanup。
Better error handling with pipeline()
pipeline() 的价值:
将整个 pipeline 作为一个整体管理,并统一处理错误和结束。
因此比裸 pipe() 更适合生产代码。
6.12 Asynchronous control flow patterns with streams
Streams 本身也可以表达:
Sequential
chunk1
→
chunk2
→
chunk3
Unordered parallel
多个 task 同时处理,完成顺序不保证。
Unordered limited parallel
限制最大并发。
Ordered parallel
虽然并行处理:
task1
task2
task3
但最终结果必须保持:
1
2
3
1
2
36.13 Piping patterns
Combining streams
多个 Transform 可以封装成一个高层 Stream。
例如:
compress
+
encrypt
=
CompressAndEncryptStream
这样可以提高复用性。
Forking streams
一个 Readable:
┌→ Writable A
Readable
└→ Writable B
适合:
- logging;
- checksum;
- audit;
- 多目的地输出。
Merging streams
多个输入:
A ─┐
B ─┼→ merged stream
C ─┘
需要处理:
- ordering;
- completion;
- backpressure。
Multiplexing and demultiplexing
Multiplex:
multiple logical streams
↓
one physical stream
Demultiplex:
one physical stream
↓
multiple logical streams
这是远程日志、网络传输等场景的重要模式。
本章核心结论
Streams 不只是”读文件的 API”,而是一套完整的增量处理、背压控制、组合式数据流架构。作者明确把 Streams 定位为处理 binary/string/object 数据的核心 Node.js pattern。
7. Creational Design Patterns
本章开始进入传统 Design Patterns,但作者特别强调:
Node.js 不应该机械照搬经典 OOP Pattern,而应该结合 JavaScript 的函数式/混合特性重新理解。
7.1 Factory
核心:
将”创建什么对象”与”调用方如何使用对象”分离。
Decoupling object creation and implementation
调用:
create(...)
↓
具体实现
而不是:
new ConcreteClass()
调用方因此不需要知道:
- class;
- constructor;
- implementation details。
A mechanism to enforce encapsulation
Factory 可以只暴露:
public object
而隐藏:
creation process
implementation
configuration
Node.js 中 Factory 可以只是一个普通 function。
7.2 Builder
解决:
constructor 参数过多、构造过程复杂。
例如:
Builder
├─ withA()
├─ withB()
├─ withC()
└─ build()
核心原则:
- 分解复杂 constructor;
- 让每一步更可读;
- builder 可以进行 validation;
- normalization;
- type conversion;
- parameter inference。
Builder 不只用于:
new Object()
也可以用于:
构造复杂函数调用。
7.3 Revealing Constructor
JavaScript 特有的重要模式。
核心思想:
在 constructor 执行阶段暴露”修改内部状态的能力”,构造完成后则不给外部这种能力。
例如:
constructor(executor)
↓
executor 获得 private modifier
↓
对象创建完成
↓
外部只有 read API
一个典型思想来源就是 Promise:
new Promise((resolve, reject) => {})
外部不能任意改变 Promise 状态。
因此这是非常强的 encapsulation。
7.4 Singleton
目的:
一个进程中共享一个实例。
常见用途:
- shared state;
- resource pool;
- database;
- configuration;
- shared service。
Node.js 中因为 module cache,Singleton 实现非常简单。
但重要 caveat:
Singleton 不是全局系统唯一
Process A → instance A
Process B → instance B
因此多进程/多机器系统中仍然可能存在多个 Singleton。
Singleton 会隐藏依赖
代码:
foo()
↓
implicitly accesses singleton
会让 dependency 不明显。
Singleton 会降低测试隔离性
共享状态可能污染 test。
因此:
Singleton 简单,但不是默认最佳方案。
7.5 Wiring modules
模块之间需要建立:
A → B → C
依赖关系。
Singleton 可以简单完成:
all → same instance
但复杂系统中 dependency graph 会隐藏。
7.6 Singleton dependencies
Singleton 的优势:
- 简单;
- 方便;
- 不需要手动传递。
缺点:
- implicit dependency;
- global-like state;
- test coupling;
- process-local。
7.7 Dependency Injection
DI:
Component
↑
dependency injected
而不是 Component 自己:
new Dependency()
核心价值:
Dependency creation 与 dependency usage 解耦。
常见形式:
Constructor injection
new Blog(db)
Function injection
foo(db)
Property injection
object.db = db
object.db = dbDI 的优势
- 可测试;
- 可替换实现;
- 可组合;
- 明确依赖;
- 减少 hidden coupling。
DI 的代价
- dependency graph 需要管理;
- 大系统 wiring 复杂;
- component 与实际 dependency 的关系不再直接可见。
因此可以使用:
- Service Locator;
- DI Container。
但这会进一步增加抽象层。
作者的总结非常明确:Factory 在 JavaScript 中非常灵活;Singleton 实现简单但有 caveat;Builder 可以用于对象和复杂函数调用;Revealing Constructor 提供强封装;Singleton 与 DI 是两种主要 module wiring 技术。
8. Structural Design Patterns
三个核心模式:
Proxy
Decorator
Adapter
原书最终给出的最重要区别:
Proxy → same interface
Decorator → enhanced interface
Adapter → different interface
8.1 Proxy
Proxy:
控制对另一个对象的访问。
Client
↓
Proxy
↓
Real Object
适合:
- logging;
- access control;
- lazy initialization;
- caching;
- remote object;
- validation;
- monitoring。
Techniques for implementing proxies
Object composition
proxy.target = target
优点:
- 简单;
- 显式;
- 可控。
Object augmentation
对对象增加/替换方法。
更灵活,但可能修改原对象语义。
Built-in Proxy object
ES6 Proxy 可以拦截:
- get;
- set;
- apply;
- construct;
- etc.
优点:
可以透明地拦截对象操作。
Change Observer with Proxy
Proxy 可以实现:
obj.foo = newValue
↓
proxy intercept
↓
emit change
从而实现 reactive/change-observer 风格。
8.2 Decorator
Decorator:
在不改变原对象核心实现的情况下,为对象增加能力。
例如:
Database
↓
LoggingDecorator
↓
CachingDecorator
↓
MetricsDecorator
Database
↓
LoggingDecorator
↓
CachingDecorator
↓
MetricsDecoratorProxy vs Decorator
两者实现技术很接近。
区别主要在意图:
Proxy
→ 控制访问
Decorator
→ 增强功能
Proxy
→ 控制访问
Decorator
→ 增强功能LevelUP plugin
典型用途:
给现有数据库 API 增加额外能力,而不修改原始实现。
8.3 Adapter
Adapter:
将已有对象的接口转换为消费者需要的接口。
Consumer
↓
Expected Interface
Adapter
↓
Existing Object
它解决的是:
接口不兼容
而不是增加功能。
8.4 Proxy / Decorator / Adapter 三者区分
| Pattern | 目标接口 | 核心目的 |
|---|---|---|
| Proxy | 相同 | 控制访问 |
| Decorator | 增强 | 增加功能 |
| Adapter | 不同 | 转换接口 |
这个判断标准非常值得记忆。
9. Behavioral Design Patterns
本章包括:
Strategy
State
Template
Iterator
Middleware
Command
原书总结强调 Strategy/State/Template 的关系,以及 Iterator、Middleware、Command 在 Node.js 中的特殊地位。
9.1 Strategy
Strategy:
将一组可互换算法封装起来。
结构:
Context
↓
Strategy
例如:
Payment
├─ CreditCardStrategy
├─ PayPalStrategy
└─ CryptoStrategy
Context 不关心具体实现。
9.2 State
State 是 Strategy 的变体。
区别:
Strategy
→ 调用方主动选择行为
State
→ 当前状态决定行为
例如:
Disconnected
Connected
Closing
Closed
每一个 state 决定对象此刻的行为。
9.3 Template
Template:
将公共流程固定,把变化部分交给子类。
结构:
Template algorithm
├─ common steps
├─ hook
└─ variable steps
可以理解为 Strategy 的”静态 OOP 版本”。
9.4 Iterator
Iterator 解决:
如何逐个访问一个集合,而不暴露其内部结构。
JavaScript 原生支持 Iterator Protocol。
核心:
next()
返回:
{
value,
done
}
{
value,
done
}Iterable protocol
对象只要实现:
Symbol.iterator
就可以:
for...of
for...ofIterators and iterables as native JS interface
这使 Iterator 不再只是经典 GoF pattern,而成为 JavaScript 的语言级能力。
9.5 Generators
Generator:
function * () {}
特点:
- 可暂停;
- 可恢复;
yield返回值;- 自动生成 iterator。
因此:
Generator 是构建 Iterator 的非常强大的语言工具。
Generators in theory
Generator 可以理解为:
一种可以在不同执行点之间来回切换的协程式控制流工具。
Controlling a generator iterator
可以:
next(value)
throw(error)
return(value)
控制 generator。
How to use generators in place of iterators
Generator 可以极大简化:
manual iterator state machine
manual iterator state machine9.6 Async iterators
Async Iterator:
next()
→ Promise
因此可以:
for await (const value of iterable)
非常适合:
- network;
- database;
- stream;
- async resource。
9.7 Async generators
:
async function * generator() {}
结合:
async
+
yield
可以非常自然地实现异步数据流。
9.8 Async iterators and Node.js streams
这是非常重要的一条连接:
Readable Stream 本质上可以看成一种异步可迭代数据源。
因此:
for await (const chunk of stream)
可以直接消费 Readable。
同时:
Readable.from(asyncIterable)
又可以把 Async Iterable 转成 Stream。
所以:
Streams ↔ Async Iterators
可以互相适配。
9.9 Middleware
Middleware 是 Node.js 生态非常典型的 Pattern。
结构:
Request
↓
Middleware 1
↓
Middleware 2
↓
Middleware 3
↓
Handler
Middleware 可以:
- preprocess;
- postprocess;
- authentication;
- logging;
- validation;
- error handling。
它本质上很接近:
Chain of Responsibility。
Middleware in Express
典型形式:
(req, res, next)
每个 middleware 决定:
continue
or
terminate
continue
or
terminateMiddleware framework
这种模式并不仅属于 HTTP。
书中用 ZeroMQ 实现 middleware framework,说明:
Middleware 是通用的 control-flow composition pattern。
9.10 Command
Command:
将”一个操作”封装成对象/值。
例如:
Command
├─ execute()
├─ undo()
├─ serialize()
└─ metadata
适合:
- undo/redo;
- queue;
- retry;
- serialization;
- scheduling;
- remote execution;
- logging;
- distributed commands。
The Task pattern
如果只是:
“把一个任务包装起来执行”
未必需要复杂 Command。
简单 Task pattern 往往足够。
因此:
简单异步任务 → Task
复杂可管理操作 → Command
简单异步任务 → Task
复杂可管理操作 → Command10. Universal JavaScript for Web Applications
目标:
同一套 JavaScript code/logic/data 在 Server 和 Browser 中尽可能复用。
原书的核心内容包括 module bundler、webpack、cross-platform branching、React、SSR、universal routing/data retrieval、two-pass rendering 和 async pages。
10.1 Sharing code with the browser
Universal JS 希望:
Server
↕
Shared code
↕
Browser
减少:
- duplicated business logic;
- duplicated validation;
- duplicated models。
10.2 JavaScript modules in a cross-platform context
问题:
Server 和 browser 能力不同。
例如:
fs → server only
DOM → browser only
因此共享代码必须:
- 明确环境边界;
- 避免直接依赖某一端 API。
10.3 Module bundlers
浏览器无法像 Node.js 一样简单地依赖大量 server-side modules,因此需要 bundler。
典型流程:
Entry
↓
Dependency Graph
↓
Resolve
↓
Transform
↓
Bundle
Entry
↓
Dependency Graph
↓
Resolve
↓
Transform
↓
BundleHow a module bundler works
核心任务:
- 找到 entry;
- 分析 import/require;
- 构建 dependency graph;
- 打包模块;
- 输出 browser-compatible bundle。
10.4 webpack
webpack 的核心不是”打包一个 JS 文件”这么简单,而是:
根据 dependency graph 进行模块解析、变换和打包。
10.5 Fundamentals of cross-platform development
核心挑战:
同一个 API
+
不同 runtime
例如:
server implementation
browser implementation
server implementation
browser implementation10.6 Runtime code branching
运行时判断:
if (typeof window !== 'undefined') ...
优点:
- 灵活;
- 运行时决定。
缺点:
- 两端代码可能一起进入 bundle;
- dead code elimination 不一定理想;
- server-only dependencies 可能进入 browser bundle。
Challenges of runtime code branching
例如:
server library
↓
browser bundle
↓
巨大 bundle / 不兼容
server library
↓
browser bundle
↓
巨大 bundle / 不兼容10.7 Build-time code branching
在 build 阶段决定:
server build
OR
browser build
优点:
- bundle 更小;
- 环境隔离更明确;
- 可以完全替换模块。
10.8 Module swapping
例如:
src/service.js
Server:
server implementation
Browser:
browser implementation
使用 bundler 在构建阶段替换。
这是跨平台架构很重要的技巧。
10.9 Design patterns for cross-platform development
核心不是某一个框架,而是:
- 共享 interface;
- 分离 platform-specific implementation;
- build-time selection;
- runtime branching;
- module swapping。
10.10 React
本章用 React 介绍组件化 UI。
核心:
Component
+
Props
+
State
=
UI
Component
+
Props
+
State
=
UIStateful components
State 驱动:
state
↓
render
↓
UI
state
↓
render
↓
UI10.11 Creating a Universal JavaScript app
目标:
Server
↓
render initial HTML
↓
Browser
↓
continue interaction
这样可以兼顾:
- 首屏速度;
- SEO;
- SPA experience。
Frontend-only app
只在 browser render:
Browser
↓
fetch data
↓
render
简单,但首屏/SEO 较弱。
10.12 Server-side rendering
SSR:
Request
↓
Server fetch data
↓
Render HTML
↓
Browser
优势:
- fast first paint;
- SEO;
- 可访问性。
10.13 Asynchronous data retrieval
SSR 的难点:
render 需要等待数据。
因此需要:
route
↓
preload data
↓
render
route
↓
preload data
↓
render10.14 Universal data retrieval
同一个页面可能:
Server:
preload data
Browser:
reuse data
避免:
server 已经请求一次
↓
browser 又请求一次
server 已经请求一次
↓
browser 又请求一次10.15 Two-pass rendering
典型:
First pass
服务器:
route
→ data
→ render
Second pass
生成/恢复客户端所需的数据上下文。
核心目标:
SSR 与 Client-side rendering 之间共享同一份初始数据。
10.16 Async pages
一个 Async Page 可以有:
data loading
loading state
error state
loaded state
并统一处理:
- SSR 已预加载;
- browser 已有 static data;
- browser 没有数据,需要请求。
本章核心结论
Universal JavaScript 的真正价值不是”Server 和 Browser 共用全部代码”,而是:
共享真正通用的部分,明确隔离 platform-specific 部分。
作者也强调具体 bundler/framework 会快速演化,但底层概念会长期存在。
11. Advanced Recipes
这一章从”通用模式”转向真实生产中常见的棘手问题。目录包括异步初始化、请求 batching/caching、取消异步操作以及 CPU-bound task。
11.1 Dealing with asynchronously initialized components
典型问题:
constructor()
↓
需要 async initialization
↓
component 尚未 ready
但 consumer 可能马上调用:
component.doSomething()
component.doSomething()Local initialization check
每个方法检查:
initialized?
简单,但问题是:
所有方法都必须重复处理初始化状态。
Delayed startup
整个系统:
initialize()
↓
ready
↓
start serving
优点:
- 简单;
- 状态清楚。
缺点:
- 某些场景无法接受整个系统 delayed startup。
Pre-initialization queues
最通用的方法:
call method
↓
not initialized
↓
queue request
↓
initialize completes
↓
replay queued requests
这样 consumer 不需要关心初始化是否完成。
本质上是:
把”尚未 ready 的调用”转换为延迟执行。
11.2 Asynchronous request batching and caching
两个非常重要的优化。
Batching
例如多个请求:
getUser(1)
getUser(2)
getUser(3)
下游支持:
getUsers([1,2,3])
则可以将多个 concurrent requests:
A
B
C
合并成:
batch(A,B,C)
减少:
- network calls;
- DB calls;
- serialization;
- overhead。
Optimal asynchronous request caching
重点不是只缓存最终结果。
还可以缓存:
正在进行中的 Promise。
例如:
Request A
↓
cache miss
↓
start Promise P
Request B
↓
same key
↓
发现 P 正在执行
↓
reuse P
这样可以避免相同请求瞬间产生大量下游调用。
本质:
cache in-flight request。
Batching + caching
两者可以结合:
cache hit
↓
return cached result
cache miss
↓
加入 batch
↓
batch execute
↓
缓存 result
↓
resolve all waiters
这是非常重要的高并发优化模式。
11.3 Canceling asynchronous operations
JavaScript 的传统 Promise 本身并不自动提供:
cancel Promise
因此需要额外设计 cancellation semantics。
A basic recipe for creating cancelable functions
核心:
operation
+
cancel()
取消后:
don't start further work
ignore result
cleanup resources
don't start further work
ignore result
cleanup resourcesWrapping asynchronous invocations
可以将普通 async function 包装成:
Cancelable operation
维护:
cancel state
cancel stateCancelable async functions with generators
Generator 可以让 cancellation 更容易控制:
yield async operation
↓
resume
↓
check cancellation
因此可以在控制流层面实现取消。
11.4 Running CPU-bound tasks
这是理解 Node.js 性能边界最重要的部分之一。
Node.js 擅长:
I/O-bound
不擅长直接在 event loop 上执行:
CPU-bound
例如:
- 大量计算;
- compression;
- image processing;
- cryptography;
- combinatorial search。
因为:
CPU task
↓
block event loop
↓
所有请求都受影响
CPU task
↓
block event loop
↓
所有请求都受影响11.5 Interleaving with setImmediate
第一种方案:
CPU task
↓
step 1
↓
setImmediate
↓
step 2
↓
setImmediate
↓
step 3
让出 event loop。
优点:
- 实现简单;
- 保持进程响应;
- 不需要额外进程。
缺点:
- 每次 yield 有 overhead;
- CPU task 总运行时间增加;
- 单个 step 太重仍然会阻塞。
process.nextTick vs setImmediate
非常重要:
不要使用
process.nextTick()来实现这种长期 CPU task interleaving。
因为 nextTick 会优先于正常 I/O 执行,递归 nextTick 可能造成:
I/O starvation
I/O starvation11.6 Using external processes
更可靠:
Main process
↓
Child process
↓
CPU task
优势:
- 主 event loop 不被阻塞;
- 可以使用多个 CPU core;
- 更容易隔离故障;
- 可以使用其他语言实现高性能任务。
适合:
重 CPU、运行时间长的任务。
Process pool
不要:
每个请求
↓
new process
进程创建本身有成本。
更好的:
Process Pool
├─ Worker 1
├─ Worker 2
└─ Worker 3
同时:
限制 worker 数量也可以防止过度资源消耗甚至 DoS 风险。
11.7 Using worker threads
Worker Threads:
Main Thread
↓
Worker Thread
与 child process 相比:
- 更轻;
- 可以共享部分内存;
- 适合 CPU-bound JS 任务;
- 不需要完整 OS process。
但也有:
- synchronization;
- memory management;
- worker lifecycle;
- serialization/transfer overhead
等复杂度。
11.8 Running CPU-bound tasks in production
选择顺序可以理解为:
任务很短
→ setImmediate interleaving
任务较重
→ Worker Threads / process pool
需要隔离 / 非 JS / 强资源隔离
→ external processes
关键不是”Node.js 不能做 CPU”,而是:
绝不能让长时间 CPU computation 占据主 event loop。
12. Scalability and Architectural Patterns
本章从 coding pattern 上升到系统架构。书中明确把 scalability 同时定义为:
- capacity;
- availability;
- failure tolerance;
- application complexity 的可扩展性。
12.1 An introduction to application scaling
Scalability:
系统随着业务、用户、数据、流量、团队规模增长仍能正常演化和运行。
不要只理解成:
QPS ↑
还包括:
complexity ↑
team size ↑
failure tolerance ↑
availability ↑
complexity ↑
team size ↑
failure tolerance ↑
availability ↑12.2 Scaling Node.js applications
Node.js 单线程模型非常适合 I/O-bound。
但单个 JavaScript thread:
最终仍有容量上限。
所以高负载时要:
1 process
↓
multiple processes
↓
multiple machines
1 process
↓
multiple processes
↓
multiple machines12.3 The three dimensions of scalability
Scale Cube:
X-axis
复制相同实例:
A
A
A
A
然后 load balance。
Y-axis
按功能拆分:
Auth service
Payment service
Inventory service
也就是:
Microservices / functional decomposition。
Z-axis
按数据/请求维度分片:
Shard A
Shard B
Shard C
即:
Data partitioning。
本书重点深入 X 和 Y。
12.4 Cloning and load balancing
最简单的扩展方式:
Load Balancer
├─ Node A
├─ Node B
└─ Node C
优点:
- 简单;
- 容错;
- 容易横向扩展。
12.5 Cluster module
Node.js cluster 可以:
Master/Primary
↓
Worker processes
使多个 process 使用同一台机器的多个 CPU core。
Notes on cluster
cluster 的本质不是:
把一个 JS thread 变成 multi-thread。
而是:
启动多个 Node.js processes。
因此每个 worker:
- 有自己的 memory;
- 有自己的 event loop;
- 有自己的 module cache。
12.6 Resiliency and availability
多个 instance 不只是为了吞吐量。
还可以:
Worker A crash
↓
Worker B/C continue
提高 availability。
12.7 Zero-downtime restart
部署:
Old workers
↓
New workers ready
↓
traffic gradually moves
↓
old workers exit
目标:
deployment 不应导致用户可见 downtime。
12.8 Dealing with stateful communications
多实例以后出现重要问题:
Request 1 → Server A
Request 2 → Server B
如果 session 在 A:
B 找不到状态
解决方案:
Sharing state
将 state 存储到:
- Redis;
- DB;
- shared storage。
Sticky load balancing
同一个 client 总是:
client → same worker
简单但降低负载均衡灵活性。
12.9 Scaling with a reverse proxy
Client
↓
Reverse Proxy
↓
Node instances
Reverse proxy 负责:
- load balancing;
- routing;
- TLS;
- connection handling;
- health checks。
Load balancing with Nginx
Nginx 是典型:
Nginx
├─ Node 1
├─ Node 2
└─ Node 3
Nginx
├─ Node 1
├─ Node 2
└─ Node 312.10 Dynamic horizontal scaling
静态配置:
server1
server2
server3
动态系统则需要:
instances add/remove
↓
traffic distribution update
因此需要 service discovery。
12.11 Service registry
Registry:
Service
↓
register
↓
Registry
↓
discover
↓
Client/Load Balancer
典型技术:
Consul
适合:
- dynamic infrastructure;
- auto scaling;
- ephemeral services。
12.12 Peer-to-peer load balancing
不一定必须:
Central LB
也可以:
Node A ↔ Node B ↔ Node C
Node 自己维护可用节点并决定请求发给谁。
优点:
- 减少中心组件;
- 更分布式。
代价:
- discovery;
- membership;
- consistency;
- failure detection
更复杂。
12.13 Scaling applications using containers
Container:
将应用及其运行环境打包成可部署单元。
价值:
- consistency;
- isolation;
- reproducibility;
- packaging;
- deployment。
12.14 Docker
Docker 的核心思想:
Image
↓
Container
把:
- application;
- dependencies;
- runtime;
- configuration
打包。
12.15 Kubernetes
Kubernetes 是 container orchestration platform。
它解决:
- deployment;
- service discovery;
- load balancing;
- scaling;
- health management;
- rollout/rollback;
- desired-state management。
其中非常重要的理念:
声明最终状态,而不是手工描述每一步操作。
12.16 Decomposing complex applications
复杂系统可以:
Monolith
↓
Decompose
↓
Services
Monolith
↓
Decompose
↓
ServicesMonolithic architecture
所有能力:
one deployment
one application
优点:
- 简单;
- 易开发;
- 易部署;
- 边界少。
缺点:
- complexity 集中;
- scaling 粗粒度;
- team coupling。
12.17 Microservice architecture
按业务能力拆:
Auth
Orders
Payments
Inventory
关键原则:
High cohesion
一个 service 内的功能高度相关。
Loose coupling
service 之间尽量减少依赖。
Microservices advantages
- 独立部署;
- 独立扩展;
- 团队自治;
- 故障隔离;
- 技术独立。
Disadvantages
最大的代价:
把代码复杂度变成了 distributed systems complexity。
新增:
- network;
- latency;
- failure;
- consistency;
- service discovery;
- deployment;
- observability;
- integration complexity。
12.18 Integration patterns
API proxy
Client
↓
API Proxy
↓
Services
用于:
- gateway;
- routing;
- auth;
- aggregation。
API orchestration
由一个 orchestrator:
Request
↓
Service A
↓
Service B
↓
Service C
↓
Aggregate result
适合复杂业务流程。
Integration with a message broker
Service A
↓
Broker
↓
Service B
优点:
- decoupling;
- async communication;
- buffering;
- reliability。
缺点:
- broker 本身需要维护;
- monitoring;
- scaling;
- operational complexity。
作者总结强调:load balancing 和 microservices 是 X/Y 两个主要扩展维度,而 microservices 并没有消除复杂度,只是把复杂度转移到了 service integration。
13. Messaging and Integration Patterns
最后一章从”如何扩展”进一步进入:
如何连接分布式系统。
书中把核心消息模型归纳为:
Publish/Subscribe
Task Distribution
Request/Reply
并分别讨论 peer-to-peer 与 broker-based 实现。
13.1 Fundamentals of a messaging system
设计 messaging system 时首先回答四个问题:
1. Direction
One-way
Request/Reply
2. Purpose
消息是什么:
Command
Event
Document
3. Timing
Synchronous
Asynchronous
4. Delivery
Peer-to-peer
Broker
这四个维度基本构成后续所有 Messaging Pattern 的坐标系。
13.2 One way versus request/reply
One-way
A → B
A 不等待返回。
适合:
- event;
- command;
- fire-and-forget。
Request/Reply
A → Request → B
A ← Reply ← B
适合:
- RPC;
- query;
- command with result。
13.3 Message types
Command Messages
表达:
请执行某件事。
例如:
CreateOrder
SendEmail
ProcessPayment
强调:
动作。
Event Messages
表达:
某件事已经发生。
例如:
OrderCreated
PaymentCompleted
PlayerLevelUp
特点:
producer 不一定关心 consumer。
因此天然适合解耦。
Document Messages
携带:
一份完整的数据描述。
例如:
{
"userId": 123,
"name": "...",
...
}
重点不是”做什么”,而是:
描述当前状态/数据。
13.4 Asynchronous messaging, queues, and streams
异步 messaging 的价值:
Producer
↓
queue/buffer
↓
Consumer
Producer 与 Consumer 不需要同时在线。
13.5 Peer-to-peer or broker-based messaging
Peer-to-peer
A ↔ B
优点:
- 少一个中心组件;
- 完全可控;
- latency 可能更低。
缺点:
- service discovery;
- reliability;
- routing;
- topology management
都由自己负责。
Broker-based
A → Broker → B
Broker 可以提供:
- buffering;
- routing;
- persistence;
- retries;
- consumer management;
- load balancing。
代价:
增加了一个必须维护和扩展的基础设施。
13.6 Publish/Subscribe
Pub/Sub:
Publisher
↓
Topic
↙ ↓ ↘
S1 S2 S3
Publisher 不需要知道具体 subscriber。
适合:
- notifications;
- events;
- chat;
- realtime updates。
13.7 Redis as a message broker
Redis Pub/Sub 可以非常简单地:
publish(topic, message)
subscribe(topic)
优点:
- 简单;
- 快;
- Node.js 集成方便。
但传统 Redis Pub/Sub 本身并不提供完整的 durable delivery guarantee。
13.8 Peer-to-peer Publish/Subscribe with ZeroMQ
ZeroMQ 提供:
- PUB;
- SUB;
以及其他 socket pattern。
它非常强调:
application 自己决定 distributed topology。
因此灵活,但需要自己承担更多架构职责。
13.9 Reliable message delivery with queues
Queue:
Producer
↓
Queue
↓
Consumer
如果 consumer 暂时不可用:
message remains queued
因此能实现可靠交付。
13.10 AMQP / RabbitMQ
AMQP 提供更完整的 messaging abstraction:
- exchanges;
- queues;
- bindings;
- acknowledgments;
- durable messages;
- consumers。
RabbitMQ 是典型实现。
Durable subscribers
重要思想:
Consumer 不在线时,消息仍可以留在 durable queue。
因此:
Consumer down
↓
messages accumulate
↓
Consumer restarts
↓
messages continue
Consumer down
↓
messages accumulate
↓
Consumer restarts
↓
messages continue13.11 Reliable messaging with streams
Stream 与 queue 不同。
Stream 更接近:
append-only log
消息有:
ordered ID
并允许:
- replay;
- history;
- multiple consumers;
- independent positions。
13.12 Streams versus message queues
Queue:
message
↓
consumer
↓
message generally consumed
Stream:
message
↓
persistent log
↓
consumer A reads
consumer B reads
consumer C reads
因此:
Queue 更强调 task delivery;Stream 更强调 durable ordered history。
13.13 Redis Streams
Redis Streams 可以用于:
- persistent messages;
- history;
- consumer groups;
- task distribution。
13.14 Task distribution patterns
目标:
Task producer
↓
Workers
↓
Results
Task producer
↓
Workers
↓
Results13.15 Fanout/Fanin
Fanout
一个任务拆成多个并行任务:
┌→ worker A
Task → ┼→ worker B
└→ worker C
Fanin
多个 worker 的结果回到:
collector
适合:
- parallel computation;
- distributed processing。
13.16 PUSH/PULL sockets
ZeroMQ:
PUSH → PULL
多个 PULL 可以形成:
load-balanced workers
重要特性:
多个 PULL consumer 会平衡收到的任务。
因此可以天然实现 distributed worker pool。
13.17 Pipelines and competing consumers
Competing Consumers:
Queue
↓
Consumer A
Consumer B
Consumer C
每条消息只交给其中一个 consumer。
适合:
task distribution。
13.18 Redis consumer groups
Redis Streams consumer groups 提供:
- 多 consumer;
- task distribution;
- consumer identity;
- processing position;
- pending entries 管理。
所以:
Stream
↓
Consumer Group
├─ Worker A
├─ Worker B
└─ Worker C
非常适合工作队列。
13.19 Request/Reply patterns
基本:
Requestor
↓
request
↓
Replier
↓
reply
↓
Requestor
最核心的问题:
当有很多 request 同时进行时,如何知道 reply 属于哪个 request?
13.20 Correlation Identifier
每一个 request:
correlationId = X
reply:
correlationId = X
Requestor:
pending[X]
收到 reply:
reply.correlationId
↓
pending[X]
↓
resolve Promise
这是实现异步 Request/Reply 的核心技术。
13.21 Return Address
Request 除了:
correlationId
还可以携带:
replyTo
即:
reply 应该发送到哪里。
于是:
Request
├─ correlationId
└─ replyTo
Replier:
send reply → replyTo
这使 Request/Reply 不必固定使用同一个响应通道。
13.22 Request/Reply + durable queue
AMQP 的一个重要结果:
Requestor
↓
durable queue
↓
Replier(s)
多个 replier:
┌→ Replier A
Queue ─┼→ Replier B
└→ Replier C
Broker 会在消费者之间分发消息,也就是 Competing Consumers。
这样:
Request/Reply 本身还可以天然获得一定的水平扩展能力。
本章核心结论
Messaging architecture 最重要的不是记 Redis、RabbitMQ、ZeroMQ 的 API,而是选择正确的 communication pattern:
事件广播
→ Publish/Subscribe
任务分发
→ Queue / Competing Consumers / PUSH-PULL
需要返回结果
→ Request/Reply
需要可靠保存
→ Queue / Stream
需要历史和 replay
→ Stream
需要最强拓扑控制
→ Peer-to-peer
需要可靠性和解耦
→ Broker
作者在结尾明确把 Publish/Subscribe、Task Distribution、Request/Reply 视为本书最重要的三类消息交换模式,并指出 Broker 可以较容易地提供可靠、可扩展的消息系统,但代价是多维护一个基础设施。
全书最值得反复复习的核心知识框架
虽然上面严格按照原书章节组织,但把 13 章串起来,可以得到一条非常清晰的学习路线:
Chapter 1
Node.js Platform
│
├─ Event Loop
├─ Reactor
├─ Non-blocking I/O
└─ libuv
│
▼
Chapter 2
Module System
│
├─ CommonJS
├─ ESM
├─ Resolution
├─ Cache
└─ Dependency Graph
│
▼
Chapter 3
Callbacks & Events
│
├─ CPS
├─ Error-first callback
├─ EventEmitter
└─ Zalgo
│
▼
Chapter 4
Async Control Flow
│
├─ Sequential
├─ Parallel
└─ Limited Parallel
│
▼
Chapter 5
Promise / Async-Await
│
├─ Promise Chain
├─ Promise.all
├─ async/await
└─ Producer-Consumer
│
▼
Chapter 6
Streams
│
├─ Streaming
├─ Backpressure
├─ Transform
├─ Pipeline
└─ Composition
│
▼
Chapter 7-9
Design Patterns
│
├─ Creational
│ ├─ Factory
│ ├─ Builder
│ ├─ Revealing Constructor
│ ├─ Singleton
│ └─ DI
│
├─ Structural
│ ├─ Proxy
│ ├─ Decorator
│ └─ Adapter
│
└─ Behavioral
├─ Strategy
├─ State
├─ Template
├─ Iterator
├─ Middleware
└─ Command
│
▼
Chapter 10
Universal JavaScript
│
├─ Bundling
├─ SSR
├─ Code sharing
└─ Universal data
│
▼
Chapter 11
Advanced Recipes
│
├─ Async initialization
├─ Batching / caching
├─ Cancellation
└─ CPU-bound work
│
▼
Chapter 12
Scalability
│
├─ X-axis cloning
├─ Y-axis decomposition
├─ Load balancing
├─ Containers
├─ Kubernetes
└─ Microservices
│
▼
Chapter 13
Messaging / Integration
│
├─ Pub/Sub
├─ Task Distribution
├─ Queue / Stream
├─ Request/Reply
├─ Correlation ID
└─ Return Address
这也是这本书真正的知识递进关系:先理解 Node.js 的运行模型,再掌握异步控制,再掌握数据流和设计模式,最后把这些能力提升到可扩展、可分布式的系统架构层面。 作者对第 6 章之后的整体定位也非常明确:Streams 是 Node.js 的关键基础模式,而后续设计模式应结合 JavaScript 的函数式/混合特性使用,而不是机械套用传统 OOP。
复习时最应该真正掌握的 20 个问题
- Node.js 为什么可以用单线程处理大量并发 I/O?
- Reactor Pattern、Event Loop、libuv 三者分别负责什么?
- 为什么 Node.js 强调 Small Core / Small Modules / Small Surface Area?
- CommonJS 和 ESM 的 module loading 机制有什么根本差异?
- CommonJS module cache 为什么会天然产生 Singleton?
- CommonJS 与 ESM 如何处理 circular dependency?
- 为什么 Zalgo 是一种严重的 API 设计问题?
- callback、EventEmitter、Promise 分别适合解决什么问题?
- Sequential / Parallel / Limited Parallel 三种异步控制流什么时候使用?
- 为什么
Array.forEach(async () => {})通常是错误的? Promise.all()为什么不能解决所有并发问题?- Backpressure 为什么是 Stream 的核心机制?
pipe()与pipeline()的区别是什么?- Proxy、Decorator、Adapter 怎么快速区分?
- Strategy、State、Template 三者是什么关系?
- 为什么 Async Iterator 与 Readable Stream 可以互相适配?
- Singleton 与 Dependency Injection 的本质 trade-off 是什么?
- Node.js 中为什么 CPU-bound task 会成为系统性能瓶颈?
- X/Y/Z Scale Cube 分别解决什么扩展问题?
- Pub/Sub、Task Distribution、Request/Reply 应该如何选择?
这 20 个问题基本覆盖了本书从运行时 → 编程模型 → 设计模式 → 系统架构 → 分布式系统的主干知识。