《Node.js Design Patterns》读书笔记


1. The Node.js Platform

本章建立后续所有章节的基础:理解 Node.js 为什么这样工作,以及为什么 Node.js 的设计方式与传统服务器端编程不同。原书将这一章的核心归结为 Node way、Reactor Pattern、V8 + libuv + Node.js Core。

1.1 The Node.js philosophy

Small core

Node.js Core 应尽可能保持小,把大量能力放到 userland。

核心思想:

  • Core 只提供稳定、基础、通用的能力。
  • 更具体的解决方案交给 npm/userland。
  • 避免核心过于庞大导致演进缓慢。
  • 让社区可以快速尝试不同方案。

思想本质:

Core 提供基础设施,生态系统提供具体解决方案。

这也是 Node.js 生态高度繁荣的重要原因。

Small modules

模块应该:

  • 小;
  • 单一职责;
  • 有明确边界;
  • 易理解;
  • 易测试;
  • 易复用。

其思想来自 Unix:

  • Small is beautiful。
  • Make each program do one thing well。

Node.js 的 npm 生态允许大量小模块共存,因此”小模块”在 Node.js 中比传统平台更加可行。

Small surface area

模块不仅应该小,而且应该拥有尽可能小的公开接口。

核心原则:

  • 不暴露内部实现。
  • 对外只提供真正需要的功能。
  • 优先暴露 function,而不是为了扩展性而暴露复杂 class hierarchy。
  • “可使用”通常比”可继承、可扩展”更重要。

结果:

API 越小:

使用方式越明确 → 错误使用越少 → 实现越简单 → 维护越容易。

Simplicity and pragmatism

Node.js 倾向于:

  • KISS;
  • “worse is better”;
  • 尽快得到一个足够好的解决方案;
  • 避免为了理论上的完美而引入巨大复杂度。

因此:

简单、可维护、实际可用,通常比理论上完美更加重要。

这也是为什么传统 GoF Pattern 在 Node.js 中经常可以被极大简化。


1.2 How Node.js works

I/O is slow

相对于 CPU 和 RAM,I/O 延迟非常高。

典型层次:

CPU/RAM → 文件系统 → 网络 → 人类输入

Node.js 的整个异步架构,本质上是在解决:

如何让 CPU 不因为等待 I/O 而闲置。

Blocking I/O

阻塞 I/O:

调用 I/O
   ↓
线程等待
   ↓
I/O 完成
   ↓
继续执行

传统解决方法:

一个连接 → 一个线程

问题:

  • 线程占用内存;
  • 上下文切换有成本;
  • 大部分时间线程实际上在等待 I/O;
  • 高并发时资源浪费严重。

Non-blocking I/O

非阻塞 I/O:

  • I/O 调用立即返回;
  • 如果暂时没有数据,则返回”暂不可用”;
  • 应用可以继续处理其他工作。

简单的 busy-wait:

循环检查资源
↓
没有数据 → 再检查
↓
有数据 → 处理

但这种方式浪费 CPU。

Event demultiplexing

更高效的方法:

Application
    ↓
Event Demultiplexer
    ↓
等待多个资源
    ↓
某个资源就绪
    ↓
返回事件
    ↓
执行对应 handler

它能够:

  • 在一个线程中管理大量 I/O;
  • 不需要轮询所有资源;
  • 只有真正有事件时才工作。

The reactor pattern

Reactor Pattern 是 Node.js 异步体系的核心。

基本模型:

Resource
   ↓
注册事件
   ↓
Reactor / Event Demultiplexer
   ↓
Event Loop
   ↓
Handler

核心流程:

  1. 注册需要监听的 I/O。
  2. Reactor 等待事件。
  3. I/O 就绪。
  4. Reactor 获取事件。
  5. 调用对应 handler。
  6. 继续等待下一批事件。

所以 Node.js 的”单线程”并不等于”一次只能处理一个请求”。

它真正的模型是:

JavaScript 执行本身是单线程的,但 I/O 并发由操作系统、libuv 和事件循环共同完成。

Libuv, the I/O engine of Node.js

Node.js 的异步能力并不是 JavaScript 自己实现的,而是建立在 libuv 上。

libuv 负责:

  • 事件循环;
  • 操作系统异步 I/O;
  • 文件系统;
  • 网络;
  • 定时器;
  • 某些无法直接异步执行的操作的线程池。

1.3 The recipe for Node.js

Node.js 可以理解为几个核心部件组合而成:

V8
 +
libuv
 +
Node.js Core APIs
 +
Bindings
 =
Node.js

V8

负责:

  • JavaScript 执行;
  • JIT;
  • Garbage Collection;
  • 内存管理。

libuv

负责底层异步 I/O 与 event loop。

Bindings

把底层能力暴露给 JavaScript。

Node.js Core Library

提供:

  • fs;
  • http;
  • stream;
  • crypto;
  • events;
  • child_process;
  • 等高级 API。

1.4 JavaScript in Node.js

Run the latest JavaScript with confidence

浏览器需要面对:

Chrome
Firefox
Safari
不同版本
不同能力

Node.js 服务端通常可以控制运行环境,因此可以:

  • 指定 Node.js 版本;
  • 使用现代 JavaScript;
  • 减少兼容性代码;
  • 减少 transpiler/polyfill 依赖。

The module system

Node.js 没有浏览器的:

  • DOM;
  • window;
  • document。

但拥有:

  • filesystem;
  • network;
  • process;
  • OS;
  • native modules。

Full access to operating system services

Node.js 的能力远大于浏览器,因为它是服务器端运行环境。

也正因为如此:

Node.js 应用的安全边界比浏览器 JavaScript 更重要。

Running native code

Node.js 可以通过:

  • Native Addons;
  • N-API;
  • WebAssembly;

调用/运行非 JavaScript 代码。


本章核心结论

Node.js = 小核心 + 小模块 + 小 API + 简单务实 + Reactor/Event Loop + V8 + libuv。

作者的总结也明确把这些作为本章核心。


2. The Module System

本章完整讨论 CommonJS 与 ESM,以及模块解析、缓存、循环依赖和互操作。原书把这两种 module system 作为 Node.js 中的两套核心模块机制。

2.1 The need for modules

模块解决的问题:

  • namespace;
  • 封装;
  • 依赖管理;
  • 复用;
  • 可测试性;
  • 大型系统拆分。

没有 module system 时:

所有代码共享 global scope
        ↓
命名冲突
        ↓
隐式依赖
        ↓
难以维护

2.2 Module systems in JavaScript and Node.js

JavaScript 主要经历:

Global scripts
   ↓
IIFE(Immediately Invoked Function Expression)
   ↓
Revealing Module Pattern
   ↓
CommonJS / AMD(Asynchronous Module Definition)
   ↓
ES Modules

Node.js 历史上主要使用 CommonJS,现代 Node.js 同时支持 ESM。


2.3 The module system and its patterns

The revealing module pattern

const myModule = (() => {
   const privateFoo = () => {}
   const privateBar = []

   const exported = {
      publicFoo: () => {},
      publicBar: () => {}
   }

   return exported
})() // once the parenthesis here are parsed, the function will be invoked

console.log(myModule)
console.log(myModule.privateFoo, myModule.privateBar)

通过 closure 隐藏 private state:

private data
    ↓
closure
    ↓
只暴露需要的 API

核心思想:

内部状态私有,对外只暴露有限接口。


2.4 CommonJS modules

A homemade module loader

function loadModule(filename, module, require) {
   const wrappedSrc =
      `(function (module, exports, require) {
      ${fs.readFileSync(filename, 'utf8')}
      })(module, module.exports, require)`
   eval(wrappedSrc)
}

function require(moduleName) {
   console.log(`Require invoked for module: ${moduleName}`)
   const id = require.resolve(moduleName)
   if (require.cached[id]) {
      return require.cached[id].exports
   }

   // module metadata
   const module = {
      exports: {},
      id
   }
   // update the cache
   require.cache[id] = module

   // load the module
   loadModule(id, module, require)

   // return exported variables
}

require.cache = {}
require.resolve = (moduleName) => {
   // resolve a full module id from the moduleName
}

理解 CommonJS 的关键是理解 require() 可以看作:

resolve
  ↓
load
  ↓
wrap
  ↓
execute
  ↓
cache
  ↓
return exports

Node.js 实际上会对模块代码进行类似 wrapper 的封装,使:

  • module;
  • exports;
  • require;
  • __filename;
  • __dirname

成为模块局部变量。


Defining a module

CommonJS:

module.exports = ...

真正导出的对象是:

module.exports

module.exports versus exports

这是最容易出错的知识点之一。

初始关系类似:

exports === module.exports

但:

exports.foo = ...

是在修改原对象。

而:

exports = ...

只是改变局部变量引用。

所以:

需要整体替换导出对象时使用 module.exports。


The require function is synchronous

CommonJS:

const foo = require('./foo')

是同步加载。

因此:

  • 模块解析发生在当前执行流程;
  • 加载/执行模块会阻塞当前 JavaScript 执行;
  • CommonJS 天然适合启动阶段加载依赖。

The resolving algorithm

require() 大致要判断:

  1. Core module
  2. File module
  3. Package module

这解释了为什么:

require('fs')
require('./foo')
require('some-package')

会走不同的解析路径。

myApp
├── foo.js
└── node_modules
    ├── depA
    │   └── index.js
    ├── depB
    │   ├── bar.js
    │   └── node_modules
    │       └── depA
    │           └── index.js
    └── depC
        ├── foobar.js
        └── node_modules
            └── depA
                └── index.js
  • Calling require('depA') from /myApp/foo.js will load /myApp/node_modules/depA/index.js
  • Calling require('depA') from /myApp/node_modules/depB/bar.js will load /myApp/node_modules/depB/node_modules/depA/index.js
  • Calling require('depA') from /myApp/node_modules/depC/foobar.js will load /myApp/node_modules/depC/node_modules/depA/index.js

The module cache

模块第一次被加载:

resolve
→ execute
→ cache

之后再次:

require()
→ cache
→ 返回同一模块实例

因此:

CommonJS module cache 天然形成”进程级共享实例”。

这也是 Node.js Singleton 实现非常简单的重要原因。

但必须注意:

Singleton / module cache 是 process-local,并不意味着整个分布式系统只有一个实例。


Circular dependencies

循环依赖:

A → B
↑   ↓
└───┘

CommonJS 的一个关键问题:

  • 模块执行具有顺序;
  • 当模块尚未完成初始化时,另一个模块可能已经拿到它的”部分 exports”。

因此可能看到:

partial exports

这也是 CommonJS circular dependency 容易出现 undefined/部分对象的原因。


2.5 Module definition patterns

Named exports

导出多个明确能力。

适合:

  • utility;
  • 多个相关函数;
  • public API。

Exporting a function

非常符合 Node.js philosophy:

一个模块可以只做一件事情。

Exporting a class

适合:

  • 有明确对象生命周期;
  • 需要实例化;
  • 维护对象状态。

但 Node.js 并不鼓励为了 OOP 而强行使用 class。

Exporting an instance

直接导出实例:

module
 ↓
instance

常用于:

  • shared state;
  • connection;
  • singleton-like component。

Modifying other modules or the global scope

包括:

  • monkey patch;
  • 修改 global;
  • 修改第三方模块。

风险很高:

  • hidden dependency;
  • 难以测试;
  • 难以推导行为;
  • 模块之间产生隐式耦合。

所以:

除非有非常明确的原因,否则避免修改其他模块或 global scope。


2.6 ESM: ECMAScript modules

Using ESM in Node.js

ESM 采用:

import ...
export ...

与 CommonJS 有明显不同。


Named exports and imports

支持:

export const foo = ...
export function bar () {}

对应:

import { foo, bar } from './module.js'

核心优点:

  • 静态结构;
  • dependency graph 可提前分析;
  • 更适合工具链;
  • tree shaking 更容易。

Default exports and imports

支持一个主要默认导出:

export default ...

适合一个模块只有一个主要概念。


Mixed exports

可以同时:

  • named exports;
  • default export。

但从 API 设计角度应避免让模块接口过于复杂。


Module identifiers

ESM 的模块标识规则与 CommonJS 不同。

重要区别:

ESM import 通常需要明确文件扩展名,而 CommonJS require() 在很多场景下可以省略。


Async imports

import() 是动态导入:

import()
   ↓
Promise
   ↓
模块加载完成

适合:

  • lazy loading;
  • 按需加载;
  • runtime 决定模块;
  • 减少初始加载成本。

2.7 Module loading in depth

Loading phases

ESM 的关键机制:

Phase 1 — Parsing / Construction

以深度优先的方式发现所有 import。

建立:

dependency graph

Phase 2 — Instantiation

创建 import/export 的绑定关系。

此时:

  • 建立引用;
  • 还没有真正执行模块代码。

Phase 3 — Evaluation

执行模块代码,给绑定赋实际值。

所以:

Parsing
→ Instantiation
→ Evaluation

这个模型是理解 ESM circular dependency 的关键。


Read-only live bindings

ESM 的 import 是:

read-only live binding

即:

  • consumer 不能重新给 imported binding 赋值;
  • 但 exporter 模块内部修改变量时,consumer 可以看到更新。

这与 CommonJS 很不同。

CommonJS 更接近:

exports object 的浅复制/对象引用语义。

ESM 更接近:

真正的符号绑定关系。


Circular dependency resolution

ESM 能比 CommonJS 更系统地处理 circular dependency,因为:

  • 解析阶段先构建完整 dependency graph;
  • instantiation 阶段建立所有 bindings;
  • evaluation 阶段再执行代码。

因此循环依赖下可以保留更加完整的引用关系。


Modifying other modules

import fs, { readFileSync } from 'fs'
import { syncBuiltinESMExports } from 'module'

fs.readFileSync = () => Buffer.from('Hello, ESM')
syncBuiltinESMExports()

console.log(fs.readFileSync === readFileSync) // true

ESM 更严格地强调:

  • import binding 是 read-only;
  • module interface 应当由模块自己定义;
  • 不应该依赖 monkey patch 式修改其他模块。

2.8 ESM and CommonJS differences and interoperability

ESM runs in strict mode

ESM 自动 strict mode。

Missing references in ESM

ESM 没有 CommonJS 的:

require
exports
module.exports
__filename
__dirname

解决方法

import { fileURLToPath } from 'url'
import { dirname } from 'path'
const __filename = fileURLToPath(import.meta.url)
const __dirname = dirname(__filename)

import { createRequire } from 'module'
const require = createRequire(import.meta.url)

this 在 ES 中是未定义的

// this.js - ESM
console.log(this) // undefined

// this.cjs - CommonJS
console.log(this === exports) // true

Interoperability

现代 Node.js 可以让 CommonJS 与 ESM 互操作,但两种模块模型本质上不同。

ESM 中导入 CommonJS(仅限默认导出)

import packageMain from 'commonjs-package' // Works
import { method } from 'commonjs-package' // Errors

ESM 中导入 JSON 数据

import { createRequire } from 'module'
const require = createRequire(import.meta.url)
const data = require('./data.json')
console.log(data)

因此设计时要清楚:

“能互相调用”不等于”两个 module system 语义完全相同”。


本章核心结论

模块系统的核心不是记住 require/import 语法,而是理解:

Module
 ├─ dependency
 ├─ encapsulation
 ├─ initialization
 ├─ caching
 ├─ resolution
 ├─ execution order
 └─ interoperability

作者最终要求掌握 CommonJS 和 ESM 两套体系,而后续书中主要采用 ESM。


3. Callbacks and Events

这一章建立 Node.js 异步编程的两个基本原语:

Callback
EventEmitter

原书称它们为 Node.js asynchronous infrastructure 的两个支柱。

3.1 The Callback pattern

Continuation-Passing Style

CPS 的思想:

函数不直接返回最终结果,而是把”接下来要执行什么”作为 callback 传进去。

普通:

result = operation()

CPS:

operation(input, callback)

Synchronous CPS

callback 立即调用:

operation()
  ↓
callback()

仍然是同步执行。


Asynchronous CPS

callback 在未来执行:

operation()
  ↓
return
  ↓
event loop
  ↓
callback()

Node.js 绝大多数 I/O API 使用这种方式。


Non-CPS callbacks

callback 不一定代表异步 continuation。

因此:

看到 callback 并不能自动推断函数是异步的。


3.2 Synchronous or asynchronous?

这是 Node.js callback API 最危险的问题之一。

一个函数可能:

某些情况同步
某些情况异步

例如:

cache hit → synchronous
cache miss → asynchronous

这种 API 会产生极难推理的问题。


An unpredictable function

调用方无法确定 callback 什么时候运行。

导致:

  • execution order 不明确;
  • stack behavior 不一致;
  • error handling 复杂;
  • 测试不稳定。

Unleashing Zalgo

Zalgo 指:

一个 API 在某些情况下同步调用 callback,在另一些情况下异步调用 callback。

这是非常危险的 API 设计。

原则:

异步 API 应该保证 callback 始终异步执行。


Using synchronous APIs

有时同步 API 本身并不是问题。

适合:

  • 启动阶段;
  • CLI;
  • 小工具;
  • 明确不需要并发的代码。

但在服务器请求处理中使用大量同步 I/O,会阻塞 event loop。


Guaranteeing asynchronicity with deferred execution

可以通过:

  • process.nextTick()
  • setImmediate()
  • 其他异步调度

把 callback 推迟执行。

但要知道:

process.nextTick() 的优先级非常高,大量递归使用可能导致 I/O starvation。


3.3 Node.js callback conventions

Node.js callback 约定:

The callback comes last

doSomething(arg1, arg2, callback)

Any error always comes first

callback(err, result)

即:

err == null
    ↓
success

否则:

error

Propagating errors

异步 callback 中:

lower-level error
       ↓
callback(err)
       ↓
upper-level callback

每一层必须正确传递 error。

典型错误:

  • 忽略 error;
  • 忘记调用 callback;
  • callback 两次;
  • error 被吞掉。

Uncaught exceptions

异步 callback 中抛出的异常不会自动像同步调用那样自然向调用方传播。

因此:

asynchronous API 中必须有清晰的 error channel。


3.4 The Observer pattern

Observer:

Subject
  ↓
notify
  ↓
Observers

Node.js 对应:

EventEmitter

The EventEmitter

核心 API:

on()
once()
emit()
removeListener()

Creating and using EventEmitter

适合:

一个操作会产生 多个事件/多个通知。

例如:

start
progress
data
error
end

Propagating errors

EventEmitter 有一个特殊事件:

error

如果没有 listener,可能导致未捕获异常。

因此:

使用 EventEmitter 时必须认真设计 error event。


Making any object observable

可以:

  • 继承 EventEmitter;
  • 组合 EventEmitter;
  • 让对象拥有事件通知能力。

重点不是继承本身,而是:

将”状态变化”与”消费者”解耦。


EventEmitter and memory leaks

listener 会持有引用。

如果:

listener 注册
↓
永远没有 remove

可能产生:

  • memory leak;
  • listener 数量增长;
  • 不必要的 callback 执行。

once() 可以自动移除一次性 listener,但如果事件永远不发生,listener 仍可能一直存在。


Synchronous and asynchronous events

同步事件:

emit()
↓
listener immediately runs

异步事件:

emit/schedule
↓
event loop
↓
listener runs

同步事件的问题是:

如果 listener 在事件产生之后才注册,就会错过事件。

因此 EventEmitter 通常更适合异步事件流。


EventEmitter versus callbacks

最重要的判断标准:

callback 用于返回一个结果;event 用于通知”发生了某件事情”。

也就是:

一次结果 → callback
多次通知 → event

Combining callbacks and events

两者可以结合:

EventEmitter → progress/events
Callback → final result

这是实际 Node.js API 中非常常见的设计。


4. Asynchronous Control Flow Patterns with Callbacks

重点从”如何异步执行一个操作”提升到:

如何组织多个异步操作。

原书特别强调 sequential、parallel、limited parallel 三类控制流。

4.1 The difficulties of asynchronous programming

异步程序最大的复杂度不是”异步”本身,而是:

多个异步任务
+
依赖关系
+
错误
+
并发
+
完成条件

Creating a simple web spider

Spider 是本章贯穿案例,用来演示:

  • recursion;
  • sequential;
  • parallel;
  • race conditions;
  • concurrency limit。

Callback hell

典型:

callback
  └─ callback
      └─ callback
          └─ callback

问题:

  • indentation;
  • readability;
  • error propagation;
  • control flow 难理解。

但真正的问题不是”嵌套太多”这么简单。

本质是:

控制流与业务逻辑混杂。


4.2 Callback best practices and control flow patterns

Callback discipline

基本原则:

一个 callback 只调用一次

避免:

success → callback()
error → callback(err)

两条路径同时执行。

尽早 return

避免:

if (...) {
  ...
} else {
  ...
}

可以:

if (err) return callback(err)

减少 nesting。

明确所有异步出口

确保:

  • success 有 callback;
  • error 有 callback;
  • empty case 有 callback;
  • early return 也有 callback。

4.3 Sequential execution

适用于:

A → B → C

B 依赖 A,C 依赖 B。


Executing a known set of tasks in sequence

如果任务集合已知:

task1
 ↓
task2
 ↓
task3

核心是维护:

current index

每完成一个任务再启动下一个。


Sequential iteration

例如:

process(items)

需要保证:

item1 完成
↓
item2
↓
item3

这适合:

  • 有顺序要求;
  • 任务之间存在资源限制;
  • 不希望瞬间产生大量并发。

4.4 Parallel execution

如果任务互不依赖:

       ┌─ Task A ─┐
Start ─┼─ Task B ─┼─> All done
       └─ Task C ─┘

性能通常更好。

核心难点:

怎么知道所有任务都结束?

经典方法:

completed counter
+
final callback

Web spider version 3

将:

A → B → C

改成:

A ─┐
B ─┼→ wait all
C ─┘

可以显著减少总耗时。


4.5 The pattern / Fixing race conditions with concurrent tasks

并行执行会产生 race condition。

例如:

Task A 修改状态
Task B 读取状态

如果顺序不确定,就可能出现错误。

因此:

并发不是简单地”同时执行所有任务”,必须明确共享状态和完成条件。


4.6 Limited parallel execution

真正生产环境中通常不能:

100,000 tasks → 同时启动

因为会导致:

  • connection exhaustion;
  • memory pressure;
  • CPU overload;
  • downstream overload。

所以需要:

max concurrency = N

Limiting concurrency

典型模型:

Queue
 ↓
N workers
 ↓
Task

保持:

activeTasks <= N

Globally limiting concurrency

有时限制不只是某一次操作,而是:

整个应用对某个资源的全局并发量。

例如:

DB max 10
HTTP max 20
file operations max 5

需要共享 concurrency controller。


4.7 The async library

本章最后强调:

实际生产环境不要轻易自己实现所有异步控制流算法。

成熟库已经解决:

  • series;
  • parallel;
  • waterfall;
  • queue;
  • retry;
  • concurrency;
  • error handling。

生产环境除非有特殊需求,否则应该优先使用成熟、经过验证的实现。


5. Asynchronous Control Flow Patterns with Promises and Async/Await

本章进入现代 JavaScript 异步核心:

Promise
+
async/await

原书强调它们能够显著简化 serial、parallel、error handling,同时注意 forEach()、return await 和递归 Promise chain 等陷阱。

5.1 Promises

What is a promise?

Promise 表示:

一个异步操作未来的结果。

三种状态:

Pending
   ↓
Fulfilled

Pending
   ↓
Rejected

一旦 settled:

fulfilled / rejected

就不能再次改变状态。


Promises/A+ and thenables

Promise/A+ 强调统一的:

then(onFulfilled, onRejected)

核心能力。

Thenable 指拥有:

then(...)

接口的对象,可以参与 Promise resolution。


5.2 The promise API

最重要:

then()
catch()
finally()

以及:

Promise.resolve()
Promise.reject()
Promise.all()
...

最关键的 Promise 特性

p2 = p1.then(...)

then() 会返回另一个 Promise。

这使得:

A
 ↓
then
 ↓
B
 ↓
then
 ↓
C

可以形成 Promise Chain。


5.3 Creating a promise

Promise executor:

new Promise((resolve, reject) => {
  ...
})

resolve

成功。

reject

失败。

重要:

executor 本身是立即执行的,而 Promise 的结果可以稍后 settle。


5.4 Promisification

把:

callback(err, result)

转换成:

Promise

例如:

old API
  ↓
promisify
  ↓
Promise API

好处:

  • 更容易组合;
  • 更容易统一 error handling;
  • 更容易使用 async/await。

5.5 Sequential execution and iteration

Promise chain 非常适合:

A
 ↓
B
 ↓
C

例如:

doA()
  .then(doB)
  .then(doC)

这就是 Promise 的经典 sequential pattern。


5.6 Parallel execution

没有依赖时:

await Promise.all([
  taskA(),
  taskB(),
  taskC()
])

即:

A ─┐
B ─┼→ Promise.all
C ─┘

优势:

  • 简洁;
  • 错误传播统一;
  • 不需要自己管理 counter。

5.7 Limited parallel execution

Promise.all() 最大的问题:

它会一次启动所有任务。

因此需要:

TaskQueue
Producer
Consumer
Concurrency Limit

这是 Producer-Consumer pattern。


5.8 Implementing the TaskQueue class with promises

典型设计:

taskQueue
consumer
concurrency

消费者数量决定并发量:

concurrency = N

队列有任务:

consumer → task

队列为空:

consumer → await

这是 Node.js 非常重要的并发控制思想。


5.9 Async/await

Async functions and await

async function:

总是返回 Promise。

await:

等待 Promise settle,并把控制权交还 event loop。

代码因此接近同步形式:

A
↓
await
↓
B
↓
await
↓
C

但底层仍然是异步的。


5.10 Error handling with async/await

核心:

try {
  await something()
} catch (err) {
  ...
}

相比 callback:

callback error forwarding

更加直观。


A unified try…catch experience

async/await 最大价值之一:

asynchronous error handling 可以重新使用熟悉的 try/catch 模型。


The “return” versus “return await” trap

两者在很多情况下看似一样:

return promise

与:

return await promise

但在 try/catch/finally 等场景中行为可能不同。

核心认识:

return await 会让当前 async function 等待 Promise settle 后再进行返回流程,因此会影响异常捕获边界。

不要机械地认为:

return await = always bad

要根据 error-handling 语义判断。


5.11 Sequential execution and iteration

最自然的方式:

for (const item of items) {
  await process(item)
}

这样天然保证:

item1 → item2 → item3

5.12 Antipattern — async/await + Array.forEach

这是全书一个非常重要的实际陷阱。

不要:

items.forEach(async item => {
  await process(item)
})

因为:

forEach() 不等待 callback 返回的 Promise。

所以外层不会等待所有任务完成。

需要 serial:

for...of + await

需要 parallel:

Promise.all(items.map(...))

5.13 Parallel execution

典型:

await Promise.all(items.map(process))

5.14 Limited parallel execution

需要:

queue
+
consumer pool

例如:

max concurrency = 10

而不是:

Promise.all(100000 tasks)

5.15 Infinite recursive promise resolution chains

一个非常容易忽略的问题:

Promise
 ↓
then
 ↓
Promise
 ↓
then
 ↓
Promise
...

大量微任务可能持续占据 event loop。

因此:

Promise 本身是异步控制工具,但不意味着无限 Promise recursion 是安全的。


6. Coding with Streams

这是 Node.js 最重要的章节之一。书中明确认为 Streams 是处理 binary data、strings、objects 的核心模式,并强调 stream composition。

6.1 Discovering the importance of streams

Buffering versus streaming

Buffering:

全部数据
 ↓
内存
 ↓
处理

Streaming:

chunk1 → process
chunk2 → process
chunk3 → process

Spatial efficiency

Streaming 最大优势:

不需要一次把全部数据放入内存。

所以特别适合:

  • 大文件;
  • HTTP;
  • 视频;
  • 数据处理;
  • 大量记录。

Time efficiency

Streaming 可以:

数据到达一点就开始处理一点。

因此:

开始处理时间

与:

整个数据读取完成时间

可以解耦。


Composability

Stream 最重要的特性之一:

Readable
  ↓
Transform
  ↓
Transform
  ↓
Writable

每个组件只负责一件事情。

这与 Node.js small modules philosophy 完美契合。


6.2 Getting started with streams

Anatomy of streams

主要四类:

Readable
Writable
Duplex
Transform

另有:

PassThrough

6.3 Readable streams

Readable 表示:

数据来源。

例如:

  • file;
  • socket;
  • stdin;
  • HTTP request。

Reading from a stream

两种主要模式:

Flowing mode

数据自动通过 data 事件流动。

Non-flowing / paused mode

消费者主动读取。


Implementing Readable streams

核心方法:

_read()

通过 push 数据进入 Readable。

重要思想:

Readable 不应该一次生成所有数据,而应该按需产生。


6.4 Writable streams

Writable 表示:

数据目的地。

例如:

  • file;
  • socket;
  • stdout;
  • database。

核心:

write(chunk)
end()

Backpressure

这是 Streams 最重要的知识点之一。

当:

Producer speed
    >
Consumer speed

就会产生:

buffer accumulation

Writable write() 的返回值就是 backpressure 信号。

大致逻辑:

write() === false
      ↓
pause producer
      ↓
wait drain
      ↓
resume

所以:

Backpressure 是生产者与消费者速度不匹配时的流量控制机制。


Implementing Writable streams

核心方法:

_write(chunk, encoding, callback)

调用 callback 表示当前 chunk 已处理。


6.5 Duplex streams

同时:

Readable
+
Writable

典型:

TCP socket

6.6 Transform streams

Transform:

input
 ↓
transform
 ↓
output

典型:

  • gzip;
  • encryption;
  • parsing;
  • filtering;
  • aggregation。

Filtering and aggregating data

Transform 不一定是一进一出:

input chunks
 ↓
filter
 ↓
aggregate
 ↓
output

因此可以做:

  • filtering;
  • grouping;
  • statistics;
  • parsing。

6.7 PassThrough streams

PassThrough:

数据基本原样通过。

常用于:

  • monitoring;
  • instrumentation;
  • debugging;
  • branching;
  • observing stream。

6.8 Observability

可以通过 PassThrough 或其他机制:

Readable
 ↓
Observable layer
 ↓
Writable

统计:

  • bytes;
  • chunks;
  • throughput;
  • timing。

6.9 Late piping

如果太晚建立 pipe,可能:

数据已经开始流动,消费者可能错过数据。

因此要理解:

  • flowing mode;
  • listener 注册时间;
  • pipe 注册时间。

6.10 Lazy streams

Stream 与 generator/iterator 配合可以实现真正 lazy 的数据来源:

不是:
Array → Stream

而是:
Generator → Stream

这样不会提前产生全部数据。


6.11 Connecting streams using pipes

基本:

readable.pipe(transform).pipe(writable)

Pipes and error handling

一个常见问题:

pipe 链中的错误并不会自动等于”所有相关资源都被安全处理”。

因此需要认真处理:

  • error;
  • close;
  • cleanup。

Better error handling with pipeline()

pipeline() 的价值:

将整个 pipeline 作为一个整体管理,并统一处理错误和结束。

因此比裸 pipe() 更适合生产代码。


6.12 Asynchronous control flow patterns with streams

Streams 本身也可以表达:

Sequential

chunk1
→
chunk2
→
chunk3

Unordered parallel

多个 task 同时处理,完成顺序不保证。

Unordered limited parallel

限制最大并发。

Ordered parallel

虽然并行处理:

task1
task2
task3

但最终结果必须保持:

1
2
3

6.13 Piping patterns

Combining streams

多个 Transform 可以封装成一个高层 Stream。

例如:

compress
+
encrypt
=
CompressAndEncryptStream

这样可以提高复用性。


Forking streams

一个 Readable:

       ┌→ Writable A
Readable
       └→ Writable B

适合:

  • logging;
  • checksum;
  • audit;
  • 多目的地输出。

Merging streams

多个输入:

A ─┐
B ─┼→ merged stream
C ─┘

需要处理:

  • ordering;
  • completion;
  • backpressure。

Multiplexing and demultiplexing

Multiplex:

multiple logical streams
       ↓
one physical stream

Demultiplex:

one physical stream
       ↓
multiple logical streams

这是远程日志、网络传输等场景的重要模式。


本章核心结论

Streams 不只是”读文件的 API”,而是一套完整的增量处理、背压控制、组合式数据流架构。作者明确把 Streams 定位为处理 binary/string/object 数据的核心 Node.js pattern。


7. Creational Design Patterns

本章开始进入传统 Design Patterns,但作者特别强调:

Node.js 不应该机械照搬经典 OOP Pattern,而应该结合 JavaScript 的函数式/混合特性重新理解。

7.1 Factory

核心:

将”创建什么对象”与”调用方如何使用对象”分离。

Decoupling object creation and implementation

调用:

create(...)
 ↓
具体实现

而不是:

new ConcreteClass()

调用方因此不需要知道:

  • class;
  • constructor;
  • implementation details。

A mechanism to enforce encapsulation

Factory 可以只暴露:

public object

而隐藏:

creation process
implementation
configuration

Node.js 中 Factory 可以只是一个普通 function。


7.2 Builder

解决:

constructor 参数过多、构造过程复杂。

例如:

Builder
 ├─ withA()
 ├─ withB()
 ├─ withC()
 └─ build()

核心原则:

  • 分解复杂 constructor;
  • 让每一步更可读;
  • builder 可以进行 validation;
  • normalization;
  • type conversion;
  • parameter inference。

Builder 不只用于:

new Object()

也可以用于:

构造复杂函数调用。


7.3 Revealing Constructor

JavaScript 特有的重要模式。

核心思想:

在 constructor 执行阶段暴露”修改内部状态的能力”,构造完成后则不给外部这种能力。

例如:

constructor(executor)
       ↓
executor 获得 private modifier
       ↓
对象创建完成
       ↓
外部只有 read API

一个典型思想来源就是 Promise:

new Promise((resolve, reject) => {})

外部不能任意改变 Promise 状态。

因此这是非常强的 encapsulation。


7.4 Singleton

目的:

一个进程中共享一个实例。

常见用途:

  • shared state;
  • resource pool;
  • database;
  • configuration;
  • shared service。

Node.js 中因为 module cache,Singleton 实现非常简单。

但重要 caveat:

Singleton 不是全局系统唯一

Process A → instance A
Process B → instance B

因此多进程/多机器系统中仍然可能存在多个 Singleton。

Singleton 会隐藏依赖

代码:

foo()
  ↓
implicitly accesses singleton

会让 dependency 不明显。

Singleton 会降低测试隔离性

共享状态可能污染 test。

因此:

Singleton 简单,但不是默认最佳方案。


7.5 Wiring modules

模块之间需要建立:

A → B → C

依赖关系。

Singleton 可以简单完成:

all → same instance

但复杂系统中 dependency graph 会隐藏。


7.6 Singleton dependencies

Singleton 的优势:

  • 简单;
  • 方便;
  • 不需要手动传递。

缺点:

  • implicit dependency;
  • global-like state;
  • test coupling;
  • process-local。

7.7 Dependency Injection

DI:

Component
   ↑
dependency injected

而不是 Component 自己:

new Dependency()

核心价值:

Dependency creation 与 dependency usage 解耦。

常见形式:

Constructor injection

new Blog(db)

Function injection

foo(db)

Property injection

object.db = db

DI 的优势

  • 可测试;
  • 可替换实现;
  • 可组合;
  • 明确依赖;
  • 减少 hidden coupling。

DI 的代价

  • dependency graph 需要管理;
  • 大系统 wiring 复杂;
  • component 与实际 dependency 的关系不再直接可见。

因此可以使用:

  • Service Locator;
  • DI Container。

但这会进一步增加抽象层。

作者的总结非常明确:Factory 在 JavaScript 中非常灵活;Singleton 实现简单但有 caveat;Builder 可以用于对象和复杂函数调用;Revealing Constructor 提供强封装;Singleton 与 DI 是两种主要 module wiring 技术。


8. Structural Design Patterns

三个核心模式:

Proxy
Decorator
Adapter

原书最终给出的最重要区别:

Proxy     → same interface
Decorator → enhanced interface
Adapter   → different interface

8.1 Proxy

Proxy:

控制对另一个对象的访问。

Client
 ↓
Proxy
 ↓
Real Object

适合:

  • logging;
  • access control;
  • lazy initialization;
  • caching;
  • remote object;
  • validation;
  • monitoring。

Techniques for implementing proxies

Object composition

proxy.target = target

优点:

  • 简单;
  • 显式;
  • 可控。

Object augmentation

对对象增加/替换方法。

更灵活,但可能修改原对象语义。

Built-in Proxy object

ES6 Proxy 可以拦截:

  • get;
  • set;
  • apply;
  • construct;
  • etc.

优点:

可以透明地拦截对象操作。


Change Observer with Proxy

Proxy 可以实现:

obj.foo = newValue
      ↓
proxy intercept
      ↓
emit change

从而实现 reactive/change-observer 风格。


8.2 Decorator

Decorator:

在不改变原对象核心实现的情况下,为对象增加能力。

例如:

Database
 ↓
LoggingDecorator
 ↓
CachingDecorator
 ↓
MetricsDecorator

Proxy vs Decorator

两者实现技术很接近。

区别主要在意图:

Proxy
→ 控制访问

Decorator
→ 增强功能

LevelUP plugin

典型用途:

给现有数据库 API 增加额外能力,而不修改原始实现。


8.3 Adapter

Adapter:

将已有对象的接口转换为消费者需要的接口。

Consumer
   ↓
Expected Interface

Adapter
   ↓

Existing Object

它解决的是:

接口不兼容

而不是增加功能。


8.4 Proxy / Decorator / Adapter 三者区分

Pattern 目标接口 核心目的
Proxy 相同 控制访问
Decorator 增强 增加功能
Adapter 不同 转换接口

这个判断标准非常值得记忆。


9. Behavioral Design Patterns

本章包括:

Strategy
State
Template
Iterator
Middleware
Command

原书总结强调 Strategy/State/Template 的关系,以及 Iterator、Middleware、Command 在 Node.js 中的特殊地位。

9.1 Strategy

Strategy:

将一组可互换算法封装起来。

结构:

Context
  ↓
Strategy

例如:

Payment
 ├─ CreditCardStrategy
 ├─ PayPalStrategy
 └─ CryptoStrategy

Context 不关心具体实现。


9.2 State

State 是 Strategy 的变体。

区别:

Strategy
→ 调用方主动选择行为

State
→ 当前状态决定行为

例如:

Disconnected
Connected
Closing
Closed

每一个 state 决定对象此刻的行为。


9.3 Template

Template:

将公共流程固定,把变化部分交给子类。

结构:

Template algorithm
 ├─ common steps
 ├─ hook
 └─ variable steps

可以理解为 Strategy 的”静态 OOP 版本”。


9.4 Iterator

Iterator 解决:

如何逐个访问一个集合,而不暴露其内部结构。

JavaScript 原生支持 Iterator Protocol。

核心:

next()

返回:

{
  value,
  done
}

Iterable protocol

对象只要实现:

Symbol.iterator

就可以:

for...of

Iterators and iterables as native JS interface

这使 Iterator 不再只是经典 GoF pattern,而成为 JavaScript 的语言级能力。


9.5 Generators

Generator:

function * () {}

特点:

  • 可暂停;
  • 可恢复;
  • yield 返回值;
  • 自动生成 iterator。

因此:

Generator 是构建 Iterator 的非常强大的语言工具。


Generators in theory

Generator 可以理解为:

一种可以在不同执行点之间来回切换的协程式控制流工具。


Controlling a generator iterator

可以:

next(value)
throw(error)
return(value)

控制 generator。


How to use generators in place of iterators

Generator 可以极大简化:

manual iterator state machine

9.6 Async iterators

Async Iterator:

next()
→ Promise

因此可以:

for await (const value of iterable)

非常适合:

  • network;
  • database;
  • stream;
  • async resource。

9.7 Async generators

:

async function * generator() {}

结合:

async
+
yield

可以非常自然地实现异步数据流。


9.8 Async iterators and Node.js streams

这是非常重要的一条连接:

Readable Stream 本质上可以看成一种异步可迭代数据源。

因此:

for await (const chunk of stream)

可以直接消费 Readable。

同时:

Readable.from(asyncIterable)

又可以把 Async Iterable 转成 Stream。

所以:

Streams ↔ Async Iterators

可以互相适配。


9.9 Middleware

Middleware 是 Node.js 生态非常典型的 Pattern。

结构:

Request
 ↓
Middleware 1
 ↓
Middleware 2
 ↓
Middleware 3
 ↓
Handler

Middleware 可以:

  • preprocess;
  • postprocess;
  • authentication;
  • logging;
  • validation;
  • error handling。

它本质上很接近:

Chain of Responsibility。


Middleware in Express

典型形式:

(req, res, next)

每个 middleware 决定:

continue
or
terminate

Middleware framework

这种模式并不仅属于 HTTP。

书中用 ZeroMQ 实现 middleware framework,说明:

Middleware 是通用的 control-flow composition pattern。


9.10 Command

Command:

将”一个操作”封装成对象/值。

例如:

Command
 ├─ execute()
 ├─ undo()
 ├─ serialize()
 └─ metadata

适合:

  • undo/redo;
  • queue;
  • retry;
  • serialization;
  • scheduling;
  • remote execution;
  • logging;
  • distributed commands。

The Task pattern

如果只是:

“把一个任务包装起来执行”

未必需要复杂 Command。

简单 Task pattern 往往足够。

因此:

简单异步任务 → Task
复杂可管理操作 → Command

10. Universal JavaScript for Web Applications

目标:

同一套 JavaScript code/logic/data 在 Server 和 Browser 中尽可能复用。

原书的核心内容包括 module bundler、webpack、cross-platform branching、React、SSR、universal routing/data retrieval、two-pass rendering 和 async pages。

10.1 Sharing code with the browser

Universal JS 希望:

Server
 ↕
Shared code
 ↕
Browser

减少:

  • duplicated business logic;
  • duplicated validation;
  • duplicated models。

10.2 JavaScript modules in a cross-platform context

问题:

Server 和 browser 能力不同。

例如:

fs → server only
DOM → browser only

因此共享代码必须:

  • 明确环境边界;
  • 避免直接依赖某一端 API。

10.3 Module bundlers

浏览器无法像 Node.js 一样简单地依赖大量 server-side modules,因此需要 bundler。

典型流程:

Entry
 ↓
Dependency Graph
 ↓
Resolve
 ↓
Transform
 ↓
Bundle

How a module bundler works

核心任务:

  1. 找到 entry;
  2. 分析 import/require;
  3. 构建 dependency graph;
  4. 打包模块;
  5. 输出 browser-compatible bundle。

10.4 webpack

webpack 的核心不是”打包一个 JS 文件”这么简单,而是:

根据 dependency graph 进行模块解析、变换和打包。


10.5 Fundamentals of cross-platform development

核心挑战:

同一个 API
+
不同 runtime

例如:

server implementation
browser implementation

10.6 Runtime code branching

运行时判断:

if (typeof window !== 'undefined') ...

优点:

  • 灵活;
  • 运行时决定。

缺点:

  • 两端代码可能一起进入 bundle;
  • dead code elimination 不一定理想;
  • server-only dependencies 可能进入 browser bundle。

Challenges of runtime code branching

例如:

server library
 ↓
browser bundle
 ↓
巨大 bundle / 不兼容

10.7 Build-time code branching

在 build 阶段决定:

server build
OR
browser build

优点:

  • bundle 更小;
  • 环境隔离更明确;
  • 可以完全替换模块。

10.8 Module swapping

例如:

src/service.js

Server:

server implementation

Browser:

browser implementation

使用 bundler 在构建阶段替换。

这是跨平台架构很重要的技巧。


10.9 Design patterns for cross-platform development

核心不是某一个框架,而是:

  • 共享 interface;
  • 分离 platform-specific implementation;
  • build-time selection;
  • runtime branching;
  • module swapping。

10.10 React

本章用 React 介绍组件化 UI。

核心:

Component
 +
Props
 +
State
 =
UI

Stateful components

State 驱动:

state
 ↓
render
 ↓
UI

10.11 Creating a Universal JavaScript app

目标:

Server
 ↓
render initial HTML
 ↓
Browser
 ↓
continue interaction

这样可以兼顾:

  • 首屏速度;
  • SEO;
  • SPA experience。

Frontend-only app

只在 browser render:

Browser
  ↓
fetch data
  ↓
render

简单,但首屏/SEO 较弱。


10.12 Server-side rendering

SSR:

Request
 ↓
Server fetch data
 ↓
Render HTML
 ↓
Browser

优势:

  • fast first paint;
  • SEO;
  • 可访问性。

10.13 Asynchronous data retrieval

SSR 的难点:

render 需要等待数据。

因此需要:

route
 ↓
preload data
 ↓
render

10.14 Universal data retrieval

同一个页面可能:

Server:
preload data

Browser:
reuse data

避免:

server 已经请求一次
↓
browser 又请求一次

10.15 Two-pass rendering

典型:

First pass

服务器:

route
→ data
→ render

Second pass

生成/恢复客户端所需的数据上下文。

核心目标:

SSR 与 Client-side rendering 之间共享同一份初始数据。


10.16 Async pages

一个 Async Page 可以有:

data loading
loading state
error state
loaded state

并统一处理:

  • SSR 已预加载;
  • browser 已有 static data;
  • browser 没有数据,需要请求。

本章核心结论

Universal JavaScript 的真正价值不是”Server 和 Browser 共用全部代码”,而是:

共享真正通用的部分,明确隔离 platform-specific 部分。

作者也强调具体 bundler/framework 会快速演化,但底层概念会长期存在。


11. Advanced Recipes

这一章从”通用模式”转向真实生产中常见的棘手问题。目录包括异步初始化、请求 batching/caching、取消异步操作以及 CPU-bound task。

11.1 Dealing with asynchronously initialized components

典型问题:

constructor()
 ↓
需要 async initialization
 ↓
component 尚未 ready

但 consumer 可能马上调用:

component.doSomething()

Local initialization check

每个方法检查:

initialized?

简单,但问题是:

所有方法都必须重复处理初始化状态。


Delayed startup

整个系统:

initialize()
 ↓
ready
 ↓
start serving

优点:

  • 简单;
  • 状态清楚。

缺点:

  • 某些场景无法接受整个系统 delayed startup。

Pre-initialization queues

最通用的方法:

call method
     ↓
not initialized
     ↓
queue request
     ↓
initialize completes
     ↓
replay queued requests

这样 consumer 不需要关心初始化是否完成。

本质上是:

把”尚未 ready 的调用”转换为延迟执行。


11.2 Asynchronous request batching and caching

两个非常重要的优化。

Batching

例如多个请求:

getUser(1)
getUser(2)
getUser(3)

下游支持:

getUsers([1,2,3])

则可以将多个 concurrent requests:

A
B
C

合并成:

batch(A,B,C)

减少:

  • network calls;
  • DB calls;
  • serialization;
  • overhead。

Optimal asynchronous request caching

重点不是只缓存最终结果。

还可以缓存:

正在进行中的 Promise。

例如:

Request A
 ↓
cache miss
 ↓
start Promise P

Request B
 ↓
same key
 ↓
发现 P 正在执行
 ↓
reuse P

这样可以避免相同请求瞬间产生大量下游调用。

本质:

cache in-flight request。


Batching + caching

两者可以结合:

cache hit
   ↓
return cached result

cache miss
   ↓
加入 batch
   ↓
batch execute
   ↓
缓存 result
   ↓
resolve all waiters

这是非常重要的高并发优化模式。


11.3 Canceling asynchronous operations

JavaScript 的传统 Promise 本身并不自动提供:

cancel Promise

因此需要额外设计 cancellation semantics。


A basic recipe for creating cancelable functions

核心:

operation
 +
cancel()

取消后:

don't start further work
ignore result
cleanup resources

Wrapping asynchronous invocations

可以将普通 async function 包装成:

Cancelable operation

维护:

cancel state

Cancelable async functions with generators

Generator 可以让 cancellation 更容易控制:

yield async operation
        ↓
resume
        ↓
check cancellation

因此可以在控制流层面实现取消。


11.4 Running CPU-bound tasks

这是理解 Node.js 性能边界最重要的部分之一。

Node.js 擅长:

I/O-bound

不擅长直接在 event loop 上执行:

CPU-bound

例如:

  • 大量计算;
  • compression;
  • image processing;
  • cryptography;
  • combinatorial search。

因为:

CPU task
 ↓
block event loop
 ↓
所有请求都受影响

11.5 Interleaving with setImmediate

第一种方案:

CPU task
 ↓
step 1
 ↓
setImmediate
 ↓
step 2
 ↓
setImmediate
 ↓
step 3

让出 event loop。

优点:

  • 实现简单;
  • 保持进程响应;
  • 不需要额外进程。

缺点:

  • 每次 yield 有 overhead;
  • CPU task 总运行时间增加;
  • 单个 step 太重仍然会阻塞。

process.nextTick vs setImmediate

非常重要:

不要使用 process.nextTick() 来实现这种长期 CPU task interleaving。

因为 nextTick 会优先于正常 I/O 执行,递归 nextTick 可能造成:

I/O starvation

11.6 Using external processes

更可靠:

Main process
      ↓
Child process
      ↓
CPU task

优势:

  • 主 event loop 不被阻塞;
  • 可以使用多个 CPU core;
  • 更容易隔离故障;
  • 可以使用其他语言实现高性能任务。

适合:

重 CPU、运行时间长的任务。


Process pool

不要:

每个请求
 ↓
new process

进程创建本身有成本。

更好的:

Process Pool
 ├─ Worker 1
 ├─ Worker 2
 └─ Worker 3

同时:

限制 worker 数量也可以防止过度资源消耗甚至 DoS 风险。


11.7 Using worker threads

Worker Threads:

Main Thread
      ↓
Worker Thread

与 child process 相比:

  • 更轻;
  • 可以共享部分内存;
  • 适合 CPU-bound JS 任务;
  • 不需要完整 OS process。

但也有:

  • synchronization;
  • memory management;
  • worker lifecycle;
  • serialization/transfer overhead

等复杂度。


11.8 Running CPU-bound tasks in production

选择顺序可以理解为:

任务很短
→ setImmediate interleaving

任务较重
→ Worker Threads / process pool

需要隔离 / 非 JS / 强资源隔离
→ external processes

关键不是”Node.js 不能做 CPU”,而是:

绝不能让长时间 CPU computation 占据主 event loop。


12. Scalability and Architectural Patterns

本章从 coding pattern 上升到系统架构。书中明确把 scalability 同时定义为:

  • capacity;
  • availability;
  • failure tolerance;
  • application complexity 的可扩展性。

12.1 An introduction to application scaling

Scalability:

系统随着业务、用户、数据、流量、团队规模增长仍能正常演化和运行。

不要只理解成:

QPS ↑

还包括:

complexity ↑
team size ↑
failure tolerance ↑
availability ↑

12.2 Scaling Node.js applications

Node.js 单线程模型非常适合 I/O-bound。

但单个 JavaScript thread:

最终仍有容量上限。

所以高负载时要:

1 process
 ↓
multiple processes
 ↓
multiple machines

12.3 The three dimensions of scalability

Scale Cube:

X-axis

复制相同实例:

A
A
A
A

然后 load balance。

Y-axis

按功能拆分:

Auth service
Payment service
Inventory service

也就是:

Microservices / functional decomposition。

Z-axis

按数据/请求维度分片:

Shard A
Shard B
Shard C

即:

Data partitioning。

本书重点深入 X 和 Y。


12.4 Cloning and load balancing

最简单的扩展方式:

Load Balancer
 ├─ Node A
 ├─ Node B
 └─ Node C

优点:

  • 简单;
  • 容错;
  • 容易横向扩展。

12.5 Cluster module

Node.js cluster 可以:

Master/Primary
       ↓
Worker processes

使多个 process 使用同一台机器的多个 CPU core。


Notes on cluster

cluster 的本质不是:

把一个 JS thread 变成 multi-thread。

而是:

启动多个 Node.js processes。

因此每个 worker:

  • 有自己的 memory;
  • 有自己的 event loop;
  • 有自己的 module cache。

12.6 Resiliency and availability

多个 instance 不只是为了吞吐量。

还可以:

Worker A crash
     ↓
Worker B/C continue

提高 availability。


12.7 Zero-downtime restart

部署:

Old workers
     ↓
New workers ready
     ↓
traffic gradually moves
     ↓
old workers exit

目标:

deployment 不应导致用户可见 downtime。


12.8 Dealing with stateful communications

多实例以后出现重要问题:

Request 1 → Server A
Request 2 → Server B

如果 session 在 A:

B 找不到状态

解决方案:

Sharing state

将 state 存储到:

  • Redis;
  • DB;
  • shared storage。

Sticky load balancing

同一个 client 总是:

client → same worker

简单但降低负载均衡灵活性。


12.9 Scaling with a reverse proxy

Client
 ↓
Reverse Proxy
 ↓
Node instances

Reverse proxy 负责:

  • load balancing;
  • routing;
  • TLS;
  • connection handling;
  • health checks。

Load balancing with Nginx

Nginx 是典型:

Nginx
 ├─ Node 1
 ├─ Node 2
 └─ Node 3

12.10 Dynamic horizontal scaling

静态配置:

server1
server2
server3

动态系统则需要:

instances add/remove
 ↓
traffic distribution update

因此需要 service discovery。


12.11 Service registry

Registry:

Service
 ↓
register
 ↓
Registry
 ↓
discover
 ↓
Client/Load Balancer

典型技术:

Consul

适合:

  • dynamic infrastructure;
  • auto scaling;
  • ephemeral services。

12.12 Peer-to-peer load balancing

不一定必须:

Central LB

也可以:

Node A ↔ Node B ↔ Node C

Node 自己维护可用节点并决定请求发给谁。

优点:

  • 减少中心组件;
  • 更分布式。

代价:

  • discovery;
  • membership;
  • consistency;
  • failure detection

更复杂。


12.13 Scaling applications using containers

Container:

将应用及其运行环境打包成可部署单元。

价值:

  • consistency;
  • isolation;
  • reproducibility;
  • packaging;
  • deployment。

12.14 Docker

Docker 的核心思想:

Image
 ↓
Container

把:

  • application;
  • dependencies;
  • runtime;
  • configuration

打包。


12.15 Kubernetes

Kubernetes 是 container orchestration platform。

它解决:

  • deployment;
  • service discovery;
  • load balancing;
  • scaling;
  • health management;
  • rollout/rollback;
  • desired-state management。

其中非常重要的理念:

声明最终状态,而不是手工描述每一步操作。


12.16 Decomposing complex applications

复杂系统可以:

Monolith
 ↓
Decompose
 ↓
Services

Monolithic architecture

所有能力:

one deployment
one application

优点:

  • 简单;
  • 易开发;
  • 易部署;
  • 边界少。

缺点:

  • complexity 集中;
  • scaling 粗粒度;
  • team coupling。

12.17 Microservice architecture

按业务能力拆:

Auth
Orders
Payments
Inventory

关键原则:

High cohesion

一个 service 内的功能高度相关。

Loose coupling

service 之间尽量减少依赖。


Microservices advantages

  • 独立部署;
  • 独立扩展;
  • 团队自治;
  • 故障隔离;
  • 技术独立。

Disadvantages

最大的代价:

把代码复杂度变成了 distributed systems complexity。

新增:

  • network;
  • latency;
  • failure;
  • consistency;
  • service discovery;
  • deployment;
  • observability;
  • integration complexity。

12.18 Integration patterns

API proxy

Client
 ↓
API Proxy
 ↓
Services

用于:

  • gateway;
  • routing;
  • auth;
  • aggregation。

API orchestration

由一个 orchestrator:

Request
 ↓
Service A
 ↓
Service B
 ↓
Service C
 ↓
Aggregate result

适合复杂业务流程。


Integration with a message broker

Service A
  ↓
Broker
  ↓
Service B

优点:

  • decoupling;
  • async communication;
  • buffering;
  • reliability。

缺点:

  • broker 本身需要维护;
  • monitoring;
  • scaling;
  • operational complexity。

作者总结强调:load balancing 和 microservices 是 X/Y 两个主要扩展维度,而 microservices 并没有消除复杂度,只是把复杂度转移到了 service integration。


13. Messaging and Integration Patterns

最后一章从”如何扩展”进一步进入:

如何连接分布式系统。

书中把核心消息模型归纳为:

Publish/Subscribe
Task Distribution
Request/Reply

并分别讨论 peer-to-peer 与 broker-based 实现。

13.1 Fundamentals of a messaging system

设计 messaging system 时首先回答四个问题:

1. Direction

One-way
Request/Reply

2. Purpose

消息是什么:

Command
Event
Document

3. Timing

Synchronous
Asynchronous

4. Delivery

Peer-to-peer
Broker

这四个维度基本构成后续所有 Messaging Pattern 的坐标系。


13.2 One way versus request/reply

One-way

A → B

A 不等待返回。

适合:

  • event;
  • command;
  • fire-and-forget。

Request/Reply

A → Request → B
A ← Reply ← B

适合:

  • RPC;
  • query;
  • command with result。

13.3 Message types

Command Messages

表达:

请执行某件事。

例如:

CreateOrder
SendEmail
ProcessPayment

强调:

动作。


Event Messages

表达:

某件事已经发生。

例如:

OrderCreated
PaymentCompleted
PlayerLevelUp

特点:

producer 不一定关心 consumer。

因此天然适合解耦。


Document Messages

携带:

一份完整的数据描述。

例如:

{
  "userId": 123,
  "name": "...",
  ...
}

重点不是”做什么”,而是:

描述当前状态/数据。


13.4 Asynchronous messaging, queues, and streams

异步 messaging 的价值:

Producer
   ↓
queue/buffer
   ↓
Consumer

Producer 与 Consumer 不需要同时在线。


13.5 Peer-to-peer or broker-based messaging

Peer-to-peer

A ↔ B

优点:

  • 少一个中心组件;
  • 完全可控;
  • latency 可能更低。

缺点:

  • service discovery;
  • reliability;
  • routing;
  • topology management

都由自己负责。


Broker-based

A → Broker → B

Broker 可以提供:

  • buffering;
  • routing;
  • persistence;
  • retries;
  • consumer management;
  • load balancing。

代价:

增加了一个必须维护和扩展的基础设施。


13.6 Publish/Subscribe

Pub/Sub:

Publisher
    ↓
 Topic
  ↙ ↓ ↘
S1 S2 S3

Publisher 不需要知道具体 subscriber。

适合:

  • notifications;
  • events;
  • chat;
  • realtime updates。

13.7 Redis as a message broker

Redis Pub/Sub 可以非常简单地:

publish(topic, message)
subscribe(topic)

优点:

  • 简单;
  • 快;
  • Node.js 集成方便。

但传统 Redis Pub/Sub 本身并不提供完整的 durable delivery guarantee。


13.8 Peer-to-peer Publish/Subscribe with ZeroMQ

ZeroMQ 提供:

  • PUB;
  • SUB;

以及其他 socket pattern。

它非常强调:

application 自己决定 distributed topology。

因此灵活,但需要自己承担更多架构职责。


13.9 Reliable message delivery with queues

Queue:

Producer
 ↓
Queue
 ↓
Consumer

如果 consumer 暂时不可用:

message remains queued

因此能实现可靠交付。


13.10 AMQP / RabbitMQ

AMQP 提供更完整的 messaging abstraction:

  • exchanges;
  • queues;
  • bindings;
  • acknowledgments;
  • durable messages;
  • consumers。

RabbitMQ 是典型实现。


Durable subscribers

重要思想:

Consumer 不在线时,消息仍可以留在 durable queue。

因此:

Consumer down
   ↓
messages accumulate
   ↓
Consumer restarts
   ↓
messages continue

13.11 Reliable messaging with streams

Stream 与 queue 不同。

Stream 更接近:

append-only log

消息有:

ordered ID

并允许:

  • replay;
  • history;
  • multiple consumers;
  • independent positions。

13.12 Streams versus message queues

Queue:

message
 ↓
consumer
 ↓
message generally consumed

Stream:

message
 ↓
persistent log
 ↓
consumer A reads
consumer B reads
consumer C reads

因此:

Queue 更强调 task delivery;Stream 更强调 durable ordered history。


13.13 Redis Streams

Redis Streams 可以用于:

  • persistent messages;
  • history;
  • consumer groups;
  • task distribution。

13.14 Task distribution patterns

目标:

Task producer
      ↓
Workers
      ↓
Results

13.15 Fanout/Fanin

Fanout

一个任务拆成多个并行任务:

       ┌→ worker A
Task → ┼→ worker B
       └→ worker C

Fanin

多个 worker 的结果回到:

collector

适合:

  • parallel computation;
  • distributed processing。

13.16 PUSH/PULL sockets

ZeroMQ:

PUSH → PULL

多个 PULL 可以形成:

load-balanced workers

重要特性:

多个 PULL consumer 会平衡收到的任务。

因此可以天然实现 distributed worker pool。


13.17 Pipelines and competing consumers

Competing Consumers:

Queue
 ↓
Consumer A
Consumer B
Consumer C

每条消息只交给其中一个 consumer。

适合:

task distribution。


13.18 Redis consumer groups

Redis Streams consumer groups 提供:

  • 多 consumer;
  • task distribution;
  • consumer identity;
  • processing position;
  • pending entries 管理。

所以:

Stream
 ↓
Consumer Group
 ├─ Worker A
 ├─ Worker B
 └─ Worker C

非常适合工作队列。


13.19 Request/Reply patterns

基本:

Requestor
   ↓
request
   ↓
Replier
   ↓
reply
   ↓
Requestor

最核心的问题:

当有很多 request 同时进行时,如何知道 reply 属于哪个 request?


13.20 Correlation Identifier

每一个 request:

correlationId = X

reply:

correlationId = X

Requestor:

pending[X]

收到 reply:

reply.correlationId
   ↓
pending[X]
   ↓
resolve Promise

这是实现异步 Request/Reply 的核心技术。


13.21 Return Address

Request 除了:

correlationId

还可以携带:

replyTo

即:

reply 应该发送到哪里。

于是:

Request
 ├─ correlationId
 └─ replyTo

Replier:

send reply → replyTo

这使 Request/Reply 不必固定使用同一个响应通道。


13.22 Request/Reply + durable queue

AMQP 的一个重要结果:

Requestor
    ↓
durable queue
    ↓
Replier(s)

多个 replier:

       ┌→ Replier A
Queue ─┼→ Replier B
       └→ Replier C

Broker 会在消费者之间分发消息,也就是 Competing Consumers。

这样:

Request/Reply 本身还可以天然获得一定的水平扩展能力。


本章核心结论

Messaging architecture 最重要的不是记 Redis、RabbitMQ、ZeroMQ 的 API,而是选择正确的 communication pattern:

事件广播
→ Publish/Subscribe

任务分发
→ Queue / Competing Consumers / PUSH-PULL

需要返回结果
→ Request/Reply

需要可靠保存
→ Queue / Stream

需要历史和 replay
→ Stream

需要最强拓扑控制
→ Peer-to-peer

需要可靠性和解耦
→ Broker

作者在结尾明确把 Publish/Subscribe、Task Distribution、Request/Reply 视为本书最重要的三类消息交换模式,并指出 Broker 可以较容易地提供可靠、可扩展的消息系统,但代价是多维护一个基础设施。


全书最值得反复复习的核心知识框架

虽然上面严格按照原书章节组织,但把 13 章串起来,可以得到一条非常清晰的学习路线:

Chapter 1
Node.js Platform
    │
    ├─ Event Loop
    ├─ Reactor
    ├─ Non-blocking I/O
    └─ libuv
          │
          ▼
Chapter 2
Module System
    │
    ├─ CommonJS
    ├─ ESM
    ├─ Resolution
    ├─ Cache
    └─ Dependency Graph
          │
          ▼
Chapter 3
Callbacks & Events
    │
    ├─ CPS
    ├─ Error-first callback
    ├─ EventEmitter
    └─ Zalgo
          │
          ▼
Chapter 4
Async Control Flow
    │
    ├─ Sequential
    ├─ Parallel
    └─ Limited Parallel
          │
          ▼
Chapter 5
Promise / Async-Await
    │
    ├─ Promise Chain
    ├─ Promise.all
    ├─ async/await
    └─ Producer-Consumer
          │
          ▼
Chapter 6
Streams
    │
    ├─ Streaming
    ├─ Backpressure
    ├─ Transform
    ├─ Pipeline
    └─ Composition
          │
          ▼
Chapter 7-9
Design Patterns
    │
    ├─ Creational
    │    ├─ Factory
    │    ├─ Builder
    │    ├─ Revealing Constructor
    │    ├─ Singleton
    │    └─ DI
    │
    ├─ Structural
    │    ├─ Proxy
    │    ├─ Decorator
    │    └─ Adapter
    │
    └─ Behavioral
         ├─ Strategy
         ├─ State
         ├─ Template
         ├─ Iterator
         ├─ Middleware
         └─ Command
          │
          ▼
Chapter 10
Universal JavaScript
    │
    ├─ Bundling
    ├─ SSR
    ├─ Code sharing
    └─ Universal data
          │
          ▼
Chapter 11
Advanced Recipes
    │
    ├─ Async initialization
    ├─ Batching / caching
    ├─ Cancellation
    └─ CPU-bound work
          │
          ▼
Chapter 12
Scalability
    │
    ├─ X-axis cloning
    ├─ Y-axis decomposition
    ├─ Load balancing
    ├─ Containers
    ├─ Kubernetes
    └─ Microservices
          │
          ▼
Chapter 13
Messaging / Integration
    │
    ├─ Pub/Sub
    ├─ Task Distribution
    ├─ Queue / Stream
    ├─ Request/Reply
    ├─ Correlation ID
    └─ Return Address

这也是这本书真正的知识递进关系:先理解 Node.js 的运行模型,再掌握异步控制,再掌握数据流和设计模式,最后把这些能力提升到可扩展、可分布式的系统架构层面。 作者对第 6 章之后的整体定位也非常明确:Streams 是 Node.js 的关键基础模式,而后续设计模式应结合 JavaScript 的函数式/混合特性使用,而不是机械套用传统 OOP。

复习时最应该真正掌握的 20 个问题

  1. Node.js 为什么可以用单线程处理大量并发 I/O?
  2. Reactor Pattern、Event Loop、libuv 三者分别负责什么?
  3. 为什么 Node.js 强调 Small Core / Small Modules / Small Surface Area?
  4. CommonJS 和 ESM 的 module loading 机制有什么根本差异?
  5. CommonJS module cache 为什么会天然产生 Singleton?
  6. CommonJS 与 ESM 如何处理 circular dependency?
  7. 为什么 Zalgo 是一种严重的 API 设计问题?
  8. callback、EventEmitter、Promise 分别适合解决什么问题?
  9. Sequential / Parallel / Limited Parallel 三种异步控制流什么时候使用?
  10. 为什么 Array.forEach(async () => {}) 通常是错误的?
  11. Promise.all() 为什么不能解决所有并发问题?
  12. Backpressure 为什么是 Stream 的核心机制?
  13. pipe() 与 pipeline() 的区别是什么?
  14. Proxy、Decorator、Adapter 怎么快速区分?
  15. Strategy、State、Template 三者是什么关系?
  16. 为什么 Async Iterator 与 Readable Stream 可以互相适配?
  17. Singleton 与 Dependency Injection 的本质 trade-off 是什么?
  18. Node.js 中为什么 CPU-bound task 会成为系统性能瓶颈?
  19. X/Y/Z Scale Cube 分别解决什么扩展问题?
  20. Pub/Sub、Task Distribution、Request/Reply 应该如何选择?

这 20 个问题基本覆盖了本书从运行时 → 编程模型 → 设计模式 → 系统架构 → 分布式系统的主干知识。


文章作者: Kiba Amor
版权声明: 本博客所有文章除特別声明外,均采用 CC BY-NC-ND 4.0 许可协议。转载请注明来源 Kiba Amor !
  目录