fix(infra): Telegram 文本报告改用 HTML 渲染并与 QQ 官方排版对齐 - #246
Conversation
Reviewer's Guide本 PR 将 Telegram 文本报告改为按需转换的 HTML 富文本,增加 HTML 失败后的纯文本兜底,消除扁平化报告的重复标题,并复用 QQ 官方 Markdown 报告排版;无提及能力时通过严格的 ID 引用识别显示用户昵称,同时兼容缺少可选接口的自定义报告生成器。 Sequence diagram for Telegram Markdown report deliverysequenceDiagram
participant Dispatcher as ReportDispatcher
participant Generator as ReportGenerator
participant Adapter as TelegramAdapter
participant Telegram as TelegramBotAPI
Dispatcher->>Generator: generate_shared_markdown_text_report(analysis_result)
Generator-->>Dispatcher: Markdown report
Dispatcher->>Adapter: send_text(group_id, text)
Adapter->>Adapter: looks_like_markdown_report(text)
alt Markdown report detected
Adapter->>Adapter: to_telegram_html(text)
Adapter->>Telegram: send_message(text, parse_mode=HTML)
alt HTML delivery fails
Telegram-->>Adapter: error
Adapter->>Adapter: strip_markdown(text)
Adapter->>Telegram: send_message(text)
end
else Ordinary text
Adapter->>Telegram: send_message(text)
end
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've found 4 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="src/infrastructure/reporting/dispatcher.py" line_range="554-560" />
<code_context>
)
+ elif self._uses_shared_markdown_report(platform_id):
+ # 无 @ 提及能力的平台复用 QQ 官方的 Markdown 排版(身份降级为昵称)
+ text_report = self.report_generator.generate_shared_markdown_text_report(
+ analysis_result
+ )
else:
</code_context>
<issue_to_address>
**issue (bug_risk):** The dispatcher calls `generate_shared_markdown_text_report` for Telegram, but that method is not part of the `IReportGenerator` contract, so a configured custom or test implementation of the interface raises `AttributeError` instead of dispatching the report.
**Triggers:** When the dispatcher is constructed with an `IReportGenerator` implementation other than the concrete `ReportGenerator` class.
**Suggested fix:** Add the method to `IReportGenerator`, or use a capability check and fall back to `generate_text_report` when the implementation does not provide it.
```suggestion
elif self._uses_shared_markdown_report(platform_id):
# 无 @ 提及能力的平台复用 QQ 官方的 Markdown 排版(身份降级为昵称)
shared_markdown_generator = getattr(
self.report_generator, "generate_shared_markdown_text_report", None
)
if callable(shared_markdown_generator):
text_report = shared_markdown_generator(analysis_result)
else:
text_report = self.report_generator.generate_text_report(analysis_result)
else:
text_report = self.report_generator.generate_text_report(analysis_result)
```
</issue_to_address>
### Comment 2
<location path="src/infrastructure/reporting/qq_official_markdown.py" line_range="198-202" />
<code_context>
+ name = names[user_id]
+ source = source.replace(f"[{user_id}]", name)
+ source = source.replace(f"<@{user_id}>", name)
+ source = re.sub(
+ rf"(?<![A-Za-z0-9_[<]){re.escape(user_id)}(?![A-Za-z0-9_>])",
+ name,
+ source,
+ )
</code_context>
<issue_to_address>
**issue (bug_risk):** The supposedly explicit-reference-only replacement also replaces any standalone occurrence of a known user ID, including an ordinary statistic or other report value. For example, if a user ID is `123`, a message count or topic text containing standalone `123` is replaced with that user's nickname.
**Triggers:** When a report value happens to equal a known user ID as a standalone token.
**Suggested fix:** Restrict replacement to explicit `[id]` and `<@id>` forms, or introduce a report-specific reference format instead of treating every standalone ID as an identity reference.
```suggestion
```
</issue_to_address>
### Comment 3
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="88-89" />
<code_context>
+ escaped = cls._MD_SUBHEADING_RE.sub(r"<b>\1</b>\n", escaped)
+ escaped = cls._MD_QUOTE_RE.sub(r"<blockquote>\1</blockquote>", escaped)
+ escaped = cls._MD_BULLET_RE.sub("• ", escaped)
+ escaped = cls._MD_BOLD_RE.sub(r"<b>\1</b>", escaped)
+ return cls._MD_CODE_RE.sub(r"<code>\1</code>", escaped)
+
+ @classmethod
</code_context>
<issue_to_address>
**nitpick (bug_risk):** Bold conversion runs before code conversion, so Markdown markers inside inline code are interpreted as HTML formatting. A literal code span such as `` `**x**` `` becomes nested `<b>` markup inside `<code>` instead of displaying the literal `**x**` content.
**Triggers:** When report content contains `**...**` inside an inline-code span.
**Suggested fix:** Protect code spans before applying bold, heading, quote, or list transformations, then restore them after the other Markdown conversions.
</issue_to_address>
### Comment 4
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="31" />
<code_context>
class TelegramMessageConverter:
"""Telegram 消息转换与昵称自愈器。"""
+ # Telegram 文本渲染:报告文本带 markdown 语法(**加粗** / *斜体* / `代码`),
+ # 而 Bot API 默认按纯文本发送,导致群内看到的是 ** 原样字符。
+ # 这里把 markdown 转成 Telegram HTML 模式可识别的标签,并转义 HTML 保留字符。
+ _MD_BOLD_RE = re.compile(r"\*\*(.+?)\*\*", re.S)
+ _MD_CODE_RE = re.compile(r"`([^`\n]+)`")
</code_context>
<issue_to_address>
**nitpick:** The class comment states that the converter handles `*斜体*`, but the implementation no longer has an italic conversion rule, so the comment documents behavior that the code does not provide.
**Suggested fix:** Remove the italic claim from the comment or document the intentional decision that italic conversion was removed.
```suggestion
# Telegram 文本渲染:报告文本带 markdown 语法(**加粗** / `代码`),
```
</issue_to_address>Sourcery assessment
Needs a human reviewer. 2 findings to address first, and this changes the automatically generated reports sent to Telegram users, including identity rendering and HTML conversion. If the conversion or name mapping is wrong, incorrect content can be sent externally and remains in those messages after a revert; reverting only fixes future reports.
Blocking findings: src/infrastructure/reporting/dispatcher.py:560, src/infrastructure/reporting/qq_official_markdown.py:202
32ed8b9 to
9c568f1
Compare
|
@sourcery-ai review |
There was a problem hiding this comment.
Hey - I've found 1 issue
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="107-108" />
<code_context>
+ @classmethod
+ def strip_markdown(cls, text: str) -> str:
+ """去除 markdown 标记,得到纯文本(HTML 渲染失败时的兜底)。"""
+ plain = cls._MD_BOLD_RE.sub(r"\1", text)
+ return cls._MD_CODE_RE.sub(r"\1", plain)
+
@staticmethod
</code_context>
<issue_to_address>
**issue (broader_impact):** 纯文本降级先执行加粗正则、再执行代码正则,因此代码段中的 `**` 会先被当成报告加粗标记处理;例如 `` `**x**` `` 会降级为 `x`,丢失代码内容中的字面量星号,而 HTML 路径已经专门通过占位符避免了这一问题。
**Triggers:** 当报告包含行内代码中的 `**`,且 HTML 发送失败时。
**Suggested fix:** 与 `to_telegram_html` 一样先摘取代码段,再处理加粗和其他 Markdown 标记,最后恢复代码内容。
</issue_to_address>Sourcery assessment
Needs a human reviewer. 1 finding to address first, and incorrect Markdown-to-HTML conversion or identity replacement could send malformed or misleading report content to Telegram users, and an HTML-send exception could potentially result in duplicate delivery if the first request succeeded. Reverting stops future messages but does not retract reports already sent outside the team.
Blocking findings: src/infrastructure/platform/adapters/telegram_message_converter.py:108
9c568f1 to
8d01c95
Compare
|
@sourcery-ai review |
Sourcery withdrew this approval because the latest commits introduced blocking findings.
There was a problem hiding this comment.
Hey - I've found 3 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="src/infrastructure/platform/adapters/telegram_adapter.py" line_range="404-413" />
<code_context>
+ message_thread_id=thread_id,
+ reply_to_message_id=reply_id,
+ )
+ except Exception as html_error:
+ logger.warning(
+ f"[Telegram] HTML 模式发送失败,降级为纯文本: {html_error}"
+ )
+ await client.send_message(
+ chat_id=chat_id,
</code_context>
<issue_to_address>
**issue (bug_risk):** When Telegram accepts the HTML request but the client receives a timeout or other transport exception, the code sends the same report again as plain text, producing duplicate report messages.
**Triggers:** When the HTML send succeeds remotely but its response is lost or times out.
**Suggested fix:** Retry only on errors that prove the request was not accepted, or use an idempotency mechanism instead of unconditionally sending the fallback.
</issue_to_address>
### Comment 2
<location path="src/infrastructure/platform/adapters/telegram_message_converter.py" line_range="74-102" />
<code_context>
+ def to_telegram_html(cls, text: str) -> str:
</code_context>
<issue_to_address>
**issue (bug_risk):** Reports longer than the adapter's 4000-character chunks are converted one chunk at a time, so Markdown delimiters split across chunks are never paired; the affected chunks retain raw `**` or backticks instead of rendering the formatting.
**Triggers:** When a long report is split inside a bold or inline-code span by `send_forward_msg`.
**Suggested fix:** Split on safe Markdown boundaries before conversion, or preserve formatting state across chunks and convert the complete report before performing a size-aware split.
</issue_to_address>
### Comment 3
<location path="src/infrastructure/reporting/qq_official_markdown.py" line_range="199-214" />
<code_context>
+ 统计数字当成身份引用(ID 为 123 的用户不会让"消息总数 123"变成昵称)。
+ """
+ source = str(text or "")
+ for user_id in sorted(names, key=len, reverse=True):
+ name = names[user_id]
+ source = source.replace(f"[{user_id}]", name)
+ source = source.replace(f"<@{user_id}>", name)
+
+ bare_ids = [
+ user_id
+ for user_id in names
+ if user_id.isdigit() and len(user_id) >= cls._MIN_BARE_ID_LENGTH
+ ]
+ for user_id in sorted(bare_ids, key=len, reverse=True):
+ source = re.sub(
+ rf"(?<![A-Za-z0-9_[<]){re.escape(user_id)}(?![A-Za-z0-9_>])",
+ names[user_id],
+ source,
+ )
+ return source.strip()
+
</code_context>
<issue_to_address>
**issue (bug_risk):** Identity replacement is performed sequentially on the already-modified source, so a nickname inserted for one user is scanned again for every other user's bare numeric ID and can be partially replaced or corrupted.
**Triggers:** When a collected nickname contains another known numeric user ID, especially in Chinese text where the regex treats digits adjacent to non-ASCII characters as standalone IDs.
**Suggested fix:** Protect explicit identity replacements with placeholders, or perform all bare-ID substitutions against the original text in one regex pass and restore the protected names afterward.
</issue_to_address>Sourcery assessment
Needs a human reviewer. 3 findings to address first, and the new Markdown-to-HTML path and identity substitution can send malformed or incorrect report content, including wrong nickname replacements or formatting, to Telegram users. Reverting prevents future messages but cannot retract reports already delivered outside the team; the impact is bounded to those sent messages.
Blocking findings: src/infrastructure/platform/adapters/telegram_adapter.py:413, src/infrastructure/platform/adapters/telegram_message_converter.py:102, src/infrastructure/reporting/qq_official_markdown.py:214
问题:报告文本自带 markdown(含转发节点 **[分析报告]** 段落),而 TelegramAdapter.send_text 调用 client.send_message 时未传 parse_mode,Bot API 按纯文本发送,群内看到 ** 等标记原样字符;同时非 QQ 官方平台走另一套纯文本模板,用户引用直接输出裸 ID(如 [8362414977]),章节也比 QQ 官方少了 「活跃时间分布」,两端体验完全不一致。 解决措施: - TelegramMessageConverter 新增 escape_html / to_telegram_html / strip_markdown / looks_like_markdown_report: **加粗**、`代码` 转 <b>/<code>,# 标题转加粗行、- 列表转 •、> 引用转 <blockquote>,并转义 & < > 三个字符; 行内代码段在 HTML 与纯文本两条路径里都先摘成占位符、其余转换做完再还原,避免 `**x**` 这类代码内容被误处理; - TelegramAdapter.send_text 改用 parse_mode="HTML" 发送,HTML 发送失败自动降级为去标记纯文本,确保不丢消息; 但仅当文本确实含 Markdown 结构(** 加粗或 # 标题行)时才走 HTML 模式 —— 图片标题、进度提示等普通文本 仍按原样纯文本发送,且不再做斜体转换,避免把正文里的 2*3*4、*.py 这类字符误渲染成标签; - 新增 QQOfficialMarkdownReportGenerator.generate_plain_markdown_report:复用 QQ 官方 Markdown 排版, 身份由 <@user_id> 降级为昵称展示;只替换 [id] / <@id> 以及长度 >= 6 的纯数字裸 ID,既不误伤「深夜」 这类含短昵称的正文,也不会把「消息总数 123」这种统计值当成用户 ID; - ReportDispatcher 对不具备 @ 提及能力的平台(当前为 Telegram)改走该排版;该方法不在 IReportGenerator 契约内,调用前做能力检查,实现缺失时回退 generate_text_report,自定义生成器不会再抛 AttributeError; - send_text 的纯文本降级只在「HTML 被 Telegram 拒绝」(BadRequest / can't parse entities)时触发; 网络/超时类错误可能已投递成功,重发会造成群内重复报告,改为直接返回失败交给上层; - 长报告分段(>4000 字符)改为按行边界切分,避免把跨行的 **加粗** / `代码` 标记从中间切断; - 转发节点扁平化不再逐段插入 **[分析报告]** 标题:平台按段落切分时所有节点同名,逐段加标题 会让群里每段前面都重复一次「[分析报告]」,现在同名节点直接拼正文,正文自带 Markdown 标题时 也不再多插一次「📊 分析报告」横幅(多来源转发等名称不同的场景仍保留逐段标题)。 效果:真机 AstrBot 4.28.1 + 真实 Telegram 群,发送后 Bot API 返回的 entities 含 4 个 blockquote 与 24 个 bold(另有自动识别的 URL / EMAIL),群内加粗、列表与引用块正常渲染,报告里的 a<b & c 正确转义 为字面量;文本报告章节与 QQ 官方一致(含活跃时间分布),用户引用显示为昵称而非裸 ID,且不再出现逐段 「[分析报告]」前缀;代码段保护、短 ID 不替换、生成器接口缺失回退三项回归点已在容器内定点复测通过 (HTML 与纯文本降级两条路径都保留代码段字面量);身份替换改为对原文单次遍历,替换出的昵称不会被二次扫描; 长文本切分保持内容无损(逐段拼接等于原文);仓库 382 项单测全过。
8d01c95 to
c9eb6d3
Compare
问题:缺少对 PR #245(Telegram 注册表群组发现与 TargetResolver 注入)、PR #246(Telegram Markdown 转 HTML 与转发排版)、PR #247(QQ 官方适配器未初始化容错)以及统一 Markdown 报告链路的独立单元测试。 解决措施:新增 test_telegram_message_converter.py 与 test_telegram_group_discovery.py,补充 QQ 官方未初始化单测,并同步更新多平台 Markdown 报告测试用例。 效果:单元测试用例从 383 个扩充至 393 个,全量执行 100% 通过,为跨平台与报告重构提供完整覆盖防护。
📝 变更说明 (Summary)
修复 Telegram 文本报告中 Markdown 标记原样显示(群里看到
分析报告)的问题,并让 Telegram 文本报告复用 QQ 官方的 Markdown 排版(身份降级为昵称),使两端观感一致。共涉及 5 个文件、+273/−21。🎯 变更背景与 STAR 描述
问题背景 (Problem)
三个叠加的症状,导致 Telegram 上的文本报告既「没渲染」又「排版走样」:
TelegramAdapter.send_text调用client.send_message时未传parse_mode,Bot API 按纯文本发送 → 群内原样显示**等标记。format_forward_nodes_to_text会把平台按段落切分的节点扁平化成一条消息;而每个节点都被命名为「分析报告」,扁平化时又逐段插入[节点名]→ 群里每段前面都重复一次「[分析报告]」。ReportDispatcher._dispatch_text只对qq_official走 QQ 官方 Markdown 模块,其它平台走旧纯文本模板 —— 同一份数据在 QQ 与 Telegram 上章节不一致(Telegram 侧连「活跃时间分布」都没有),用户引用还会直接输出裸 ID(如[8362414977])。另外自查时发现
send_text是所有文本的公共出口(图片标题、进度提示、命令回复都走它),若无差别做 Markdown 转换会误伤普通文本(详见「补充说明」)。版本溯源:v5.6.3 的
send_text同样只传chat_id/text,属既有问题、并非本次重构引入。解决措施 (Action)
TelegramMessageConverter新增escape_html/to_telegram_html/strip_markdown/looks_like_markdown_report:加粗、代码转/,# 标题转加粗行、- 列表转•、> 引用转TelegramAdapter.send_text改用parse_mode=「HTML」发送,HTML 失败时自动降级为strip_markdown纯文本,保证不丢消息。同时加了硬门控:仅当文本确实含 Markdown 结构(**或# 标题行)时才走 HTML 模式,其余文本按原样纯文本发送。format_forward_nodes_to_text:同名节点不再逐段插入标题,正文自带 Markdown 标题时也不再补插「📊 分析报告」横幅;名称确实不同(典型的多来源转发)时仍保留逐段标题。QQOfficialMarkdownReportGenerator.generate_plain_markdown_report:复用 QQ 官方 Markdown 排版,身份由<@user_id>降级为昵称,且只替换明确的 ID 引用 ——[id]、<@id>以及长度 ≥ 6 的纯数字裸 ID(平台用户 ID 的长度下限),不做昵称子串替换,也不会把「消息总数 123」这类统计值当成身份引用。ReportDispatcher通过_SHARED_MARKDOWN_PLATFORMS(当前为{「telegram」})决定是否走该排版;由于该方法不在IReportGenerator契约内,调用前做能力检查,实现缺失时回退generate_text_report,自定义生成器不会抛AttributeError。验证效果 (Result)
Message.entities含 24 个bold与 4 个blockquote(消息 id 23;另有自动识别的 1 个 URL 与 1 个 EMAIL),群内加粗、•列表、引用块均正常渲染。a<b & c正确转义为字面量,未触发 HTML 解析错误。pytest tests/382 passed;ruff check/ruff format --check无告警;pyright与未打补丁的 v5.6.5 基线逐条比对为 0 新增诊断。🔗 关联 Issue (Related Issue)
📦 变更类型 (Type of Change)
feat: 新增功能 / 新增 WebUI 模块fix: 修复 Bug / 异常处理docs: 文档或注释更新style/ ♻️refactor: 报告排版与渲染路径统一perf: 性能优化test: 增加或修改单元测试 / 冒烟测试chore/ 👷ci: 构建流程 / 工作流 / 依赖包版本变动📋 提交前自检清单 (Checklist)
infra,正文含中文「问题/解决措施/效果」三点论)dashboard/与pages/)ruff check/ruff format --check无告警,pytest tests/382 passedruff check、ruff format --check、npx pyright、node scripts/verify-commit.js),结果全部通过📌 补充说明
回归自查(重要):
send_text是所有文本(含图片标题、进度提示)的公共出口,因此逐条核对过「非报告文本」不被误改。实测中抓到两种真实误伤并已消除:234=24→234=24C:/logs/.txt与.py→ 被当作一对斜体跨段吞掉处理方式:彻底移除斜体转换(报告本身只用
**加粗,收益低误伤高),并加上looks_like_markdown_report硬门控。图片标题、进度提示、数学式子、通配符路径、-/>开头的普通文本现均保持原样纯文本发送,行为与改动前逐字一致。payload 兼容性:QQ 生成器多读了
topic.contributor_ids,该字段在SummaryTopic中有default_factory=list、反序列化走.get(「contributor_ids」, [])→ 老 checkpoint 不会AttributeError。QQ 官方路径行为不变:
mention_style默认为qq,仅 Telegram 走昵称降级分支。与其它 PR 的关系:本 PR 与「群列表恒空」修复都改
telegram_adapter.py,但落点分别在约 118 行与约 380 行,互不重叠,任意顺序合并均可(已验证两两合并与三合一合并均无冲突)。顺带发现(本 PR 未修改,建议单独修):
qq_official_markdown.render_identity_text在把昵称替换为 @提及 时使用裸子串替换source.replace(name, replacement)—— 群内存在单字昵称(如「夜」)时会把「深夜/熬夜/夜间」全部插上 @;同一昵称对应多个 ID 时替换值为空字符串,会把昵称整个删掉。该函数对用户 ID 的替换有边界保护,唯独昵称这一行没有。Telegram 不经过这段代码,故未动。🔍 评审回应(Sourcery 首轮)
四条意见逐条处理,代码均已改动并复测:
generate_shared_markdown_text_report不在IReportGenerator契约内,自定义实现会抛AttributeError—— 属实,已按建议改为getattr能力检查 + 回退generate_text_report。补充一点供参考:紧邻的 QQ 官方分支调用的
generate_qq_official_markdown_report同样不在该契约内,属仓库既有写法,本 PR 未动它(避免扩大改动面)。123时「消息总数 123」会被替换成昵称) —— 属实,已收窄为只替换[id]、<@id>与长度 ≥ 6 的纯数字裸 ID,并加了常量_MIN_BARE_ID_LENGTH与注释说明(QQ 号 6~11 位、Telegram ≤10 位,6 位下限即为此设)。x会被当成嵌套加粗—— 属实,已改为先把代码段摘成占位符、其余转换做完再还原;实测x现在输出x。斜体,但实现已无斜体转换 —— 属实(斜体正是本 PR 主动移除的,用于避免234、*.py被误渲染),注释已删掉斜体说法并注明移除原因。复测证据:容器内对四个点逐项定点验证通过(代码段保护、短 ID 不替换、长 ID /
[id]/<@id>正常替换、注册表接口缺失打 ERROR 并回退),真机真实群文本报告 entities 仍为 24 bold + 4 blockquote,仓库 382 项单测全过。Summary by Sourcery
统一 Telegram 文本报告的富文本渲染与 QQ 官方排版,并增强发送失败降级和兼容性处理。
Bug Fixes:
Enhancements:
Chores: