联网搜索能力集成: opencode webfetch → 采集器

新增:
- scripts/opencode_search.py: 通过 npx opencode run 调用 webfetch 联网搜索
- 搜索缓存每日 02:30 自动刷新 (scheduler)
- 管理后台「搜索缓存」模块 + 立即运行按钮
- POST /api/system/refresh-search-cache/run 手动触发端点
- 8个分类搜索词从sources.yaml读取,调用AI联网搜索真实内容

机制:
  Python脚本 → npx opencode run → AI webfetch → 真实搜索结果 → 写入search_cache.json
  → 采集器读缓存 → LLM基于真实数据生成选题

不再需要API Key,不依赖任何搜索引擎,搜索结果来自AI的webfetch能力
This commit is contained in:
Yuzhiran Dev
2026-05-21 09:01:02 +08:00
parent 7d50878c7d
commit 6f81167691
5 changed files with 256 additions and 47 deletions
+52 -46
View File
@@ -101,74 +101,80 @@
],
"碳账户 碳普惠 个人碳减排 蚂蚁森林 2026": [
{
"title": "武汉200万市民开通碳账户,碳普惠平台加速落地",
"url": "https://36kr.com/p/3816942423000196",
"content": "个人碳账户制度在多城市试点推行,绿色消费可兑换积分权益",
"source": "webfetch"
"title": "IIGF观点 | 浅析我国碳账户体系发展现状及未来展望",
"url": "https://iigf.cufe.edu.cn/info/1012/6225.htm",
"content": "2022年我国碳账户探索步入活跃期。2016年支付宝上线'蚂蚁森林'个人碳账户,用户通过低碳行为收集虚拟能量,能量积攒到一定数额可通过公益组织在现实世界种树,截至2020年5月参与者已达5.5亿。2022年十余家机构先后推出个人碳账户,将碳账户与普惠金融挂钩——积分越多信用等级越高,可享信贷利率优惠。浙江衢州已建立覆盖工业、农业、能源等七大领域的239.6万个碳账户,发放企业碳账户贷款294亿元、个人碳账户贷款48亿元。碳账户面临三大问题:碳排放数据采集核算缺乏统一标准、各平台数据无互联互通机制、数据安全缺乏法规保障。",
"source": "opencode_webfetch"
},
{
"title": "蚂蚁森林用户突破8亿,个人碳减排量超3000万吨",
"url": "https://36kr.com/p/3818280989443202",
"content": "数字平台推动公众参与低碳生活,蚂蚁森林成为全球最大个人碳减排项目",
"source": "webfetch"
"title": "中信碳账户3周年:共创碳普惠行业标准 焕发可持续消费潜能",
"url": "https://www.citicbank.com/about/companynews/banknew/message/202504/t20250423_3563037.html",
"content": "2025年4月22日第56个世界地球日,国内首个银行主导的个人碳账户'中信碳账户'上线3周年,用户超2150万,累计碳减排量超19万吨。中信信用卡联合中汇信碳推出全国首个银行业无纸化金融场景碳普惠方法学,对电子借记卡、电子信用卡、电子账单、线上缴费、线上贷款等7类场景碳减排量进行科学量化并计入碳账户。全新'绿信分'体系从四大维度记录用户绿色生活足迹。2024年'绿色消费'主题活动吸引超4000万人次参与。深圳自2021年印发碳普惠体系建设方案以来已发布6份碳普惠方法学,形成完整制度体系。",
"source": "opencode_webfetch"
},
{
"title": "碳交易市场扩容:个人碳资产如何变现",
"url": "https://36kr.com/p/3817224182825862",
"content": "随着全国碳市场扩容,个人碳减排量的价值正在被发现",
"source": "webfetch"
"title": "个人碳账户助推绿色新风尚 — 新华网",
"url": "https://www.xinhuanet.com/fortune/20240722/c928f68f99bc4416bd0027da30203bce/c.html",
"content": "生态环境部宣教中心与中华环保联合会发布的《中国碳普惠发展与实践进展报告(2023)》显示,我国碳普惠取得显著进展。垃圾分类、绿色出行、光盘行动等日常行为均可被'个人碳账户'记录并换取收益。四川泸州'绿芽积分'小程序注册用户超35万,日活超4万,累计减碳超320吨。'个人数字碳账本'已服务北京'绿色生活季'、山西'三晋绿色生活'、黑龙江'碳惠冰城'等平台。全国碳普惠平台和碳账户产品已达上百种,推出主体包括地方政府、互联网平台企业(有ESG需求)、金融机构三类。专家建议加快出台碳普惠顶层设计政策,统一标准。",
"source": "opencode_webfetch"
},
{
"title": "中国个人碳账户前路在何方? — 对话地球",
"url": "https://dialogue.earth/zh/3/60096048/",
"content": "过去三年中国经历了全球最大规模的个人碳足迹核算实践。2015年广州启动国内最早的个人碳账户,2016年支付宝'蚂蚁森林'上线,用户通过低碳行为获'绿色能量'折算成现实树木。2022年碳账户产品从不到10个猛增到60多个,至少7家科技企业(腾讯、美团、阿里)、7家银行、16个城市和4个省份参与。2024年中国自愿碳市场将个人碳账户平台排除在外,碳普惠未能被纳入全国碳市场,也未出台相关法规推广。武汉市政府运营的碳账户允许用户用积分抵扣房贷利息(4.5万克碳积分抵扣90元贷款),低碳行为计量标准为:公交每次212.5克、地铁每公里78.4克、骑行每公里93.3克碳减排量。",
"source": "opencode_webfetch"
}
],
"环保科技 绿色产品 可持续材料 2026": [
{
"title": "环保科技崛起:可持续材料市场规模突破万亿",
"url": "https://36kr.com/p/3816060007751430",
"content": "生物基材料、可降解塑料、再生材料在消费领域的应用加速",
"source": "webfetch"
"title": "《中国消费市场绿色低碳趋势调查报告(2025—2026)》重磅发布",
"url": "https://news.qq.com/rain/a/20260422A01E4A00",
"content": "报告揭示了消费市场从'单点减碳'到'全链共生'、从'概念营销'到'价值共创'的转型趋势。2026年4月22日发布,涵盖绿色消费理念在供应链各环节的渗透与落地实践。",
"source": "opencode_webfetch"
},
{
"title": "绿色产品认证体系完善,消费者愿意为环保支付溢价",
"url": "https://36kr.com/p/3817849322964098",
"content": "第三方绿色认证逐步建立信任机制,可持续消费成为主流",
"source": "webfetch"
"title": "2026 可持续包装趋势指南:企业必看的环保包装新方向",
"url": "https://www.dhl.com/discover/zh-cn/logistics-advice/sustainability-and-green-logistics/sustainable-packaging-trends",
"content": "DHL发布,指出2026年可持续包装六大趋势:新一代生物可降解材料(PLA、菌丝体基包装)爆发增长;循环包装模式(押金返还机制)加速普及;智能包装借助二维码/NFC引导回收;轻量化设计降低运输碳排放;个性化环保轻奢风兴起;全球一次性塑料禁令与回收含量法规持续趋严。",
"source": "opencode_webfetch"
},
{
"title": "科技与自然的融合:2026年最具创新性的环保材料",
"url": "https://sspai.com/post/109743",
"content": "从蘑菇皮革到藻类生物塑料,环保科技正在改变我们的生活方式",
"source": "webfetch"
"title": "工业产品绿色设计指南(2026年版)",
"url": "http://www.ecopv.org.cn/upload/gfzwh/file/20260428/1777343910650427.pdf",
"content": "国家层面发布的官方指南,涵盖长寿命、无害化、轻量化、节能、节水、节材、降噪、节空间、易回收再生、可重复使用、零碳等11大绿色设计重点方向。针对汽车、工程机械、风电、光伏、锂电池、家用电器、纺织等行业提出具体解决方案,推动'人工智能+绿色设计'及标准体系建设。",
"source": "opencode_webfetch"
},
{
"title": "2025-2026年中国绿色消费行为白皮书 - 艾媒咨询",
"url": "https://www.sohu.com/a/997098772_121864818",
"content": "报告显示绿色消费核心受众为21-40岁中青年群体(占比78.69%),已婚已育人群达66.70%,家庭育儿需求是重要驱动因素。健康意识与环保责任是消费者购买绿色产品两大核心动因,消费场景呈现'生存型>生活型>享受型'递减趋势。",
"source": "opencode_webfetch"
}
],
"AI工具 人工智能 效率提升 2026": [
{
"title": "DeepSeek发布新一代推理模型,编程能力超越GPT-4",
"url": "https://www.deepseek.com",
"content": "深度求索发布全新AI模型,在编程和推理任务上达到国际领先水平",
"source": "webfetch"
"title": "2026年必备的40个AI工具软件,办公效率提升120%",
"url": "https://zhuanlan.zhihu.com/p/2009728985890828767",
"content": "知乎专栏文章,系统梳理2026年最值得关注的40款AI工具,涵盖通用大模型(ChatGPT/Claude/Gemini/DeepSeek)、AI思维导图(boardmix/Miro)、AI编程(GitHub Copilot/Cursor)、AI写作(Notion AI/Grammarly/Jasper)、AI绘图(Midjourney/Stable Diffusion/Adobe Firefly)、AI视频(Runway/Synthesia/CapCut)、AI音频(ElevenLabs/Suno)、AI生成PPT(博思AIPPT/Gamma)八大类,每类均分析核心优势与局限性,并提供选型指南。",
"source": "opencode_webfetch"
},
{
"title": "AI工具集网站收录AI应用超5000个,中国AI工具市场爆发",
"url": "https://ai-bot.cn/favorites/websites-to-learn-ai",
"content": "从AI写作到AI编程,中国AI工具生态正在快速成熟",
"source": "webfetch"
"title": "2026年真正实用的10款AI神器:让创作、效率与思维整理实现质的飞跃",
"url": "https://blog.csdn.net/lgf228/article/details/157800617",
"content": "CSDN技术博客精选10款职场AI工具,涵盖对话助手(DeepSeek免费全能/通义千问阿里生态/豆包内容加速)、办公效率(WPS AI深度集成/ChatExcel自然语言操作表格/飞书AI会议纪要自动化)、内容创作(Midjourney/即梦本土绘图/可灵AI视频)、思维整理(博思白板AI生成思维导图/GetNote知识管理)。包含行业实践案例:锡盟融媒体中心用DeepSeek+即梦将内容生产效率提升30%+,影视制作团队从4-6人缩减至1-2人。",
"source": "opencode_webfetch"
},
{
"title": "阿里巴巴发布Qwen3.7-Max,国产大模型综合能力登顶",
"url": "https://36kr.com/p/3818280989443202",
"content": "Qwen3.7-Max在Arena全球大模型盲测总榜中位列国产模型第一",
"source": "webfetch"
"title": "2026年成熟企业提升工作流程效率的顶级人工智能工具",
"url": "https://www.ranktracker.com/zh/blog/top-ai-tools-business-workflow-efficiency/",
"content": "RankTracker企业效率专题,分析AI在零售(Yieldigo AI定价优化)、流程自动化(UiPath RPA机器人)、品牌设计(Design.com AI标志生成)、知识管理(Notion AI)四大场景的落地实践。指出到2026年AI已从实验性技术变为企业日常工具,企业正用AI自动化重复任务、加速数据分析、优化客户支持与营销活动、支持财务预测和产品开发。",
"source": "opencode_webfetch"
},
{
"title": "Google Gemini月活用户达9亿,AI应用进入爆发期",
"url": "https://36kr.com/p/3818280989443202",
"content": "Gemini日请求量增长超7倍,AI正在渗透到日常生活的方方面面",
"source": "webfetch"
},
{
"title": "OpenAI将上市,估值高达1万亿美元",
"url": "https://36kr.com/p/3818280989443202",
"content": "OpenAI准备提交IPO文件,计划秋季上市,估值可能高达约1万亿美元",
"source": "webfetch"
"title": "2026年AI技术演进与职场变革全景展望:探索效率突破与职业转型新路径",
"url": "https://zhuanlan.zhihu.com/p/1981743817511178354",
"content": "知乎深度分析文章,指出开源框架(如DeepSeek-MoE)与轻量化模型(Phi-3-mini等38亿参数模型)持续降低AI开发门槛;企业微信AI机器人、微信'元宝'助手、支付宝智能客服等应用将AI无缝嵌入日常生产生活场景,推动'人工智能+千行百业'进入规模化落地新阶段。探讨AI对职场变革的双向影响:效率突破与职业转型路径。",
"source": "opencode_webfetch"
}
]
}
+18
View File
@@ -3,6 +3,7 @@ from fastapi import APIRouter, HTTPException, Depends, Body
from sqlalchemy.orm import Session
from sqlalchemy import func
from datetime import datetime, date
from pathlib import Path
from typing import Dict, Any, List, Optional
from pathlib import Path
import os
@@ -176,6 +177,22 @@ def trigger_metrics_sync():
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
@router.post("/refresh-search-cache/run")
def trigger_refresh_search_cache():
"""手动刷新搜索缓存"""
try:
import subprocess, sys as sys_mod
scripts_dir = Path(__file__).parent.parent.parent.parent / "scripts"
result = subprocess.run(
[sys_mod.executable, str(scripts_dir / "opencode_search.py"), "--refresh-cache"],
capture_output=True, text=True, timeout=600
)
if result.returncode != 0:
raise Exception(result.stderr[-500:])
return {"message": "搜索缓存已刷新", "output": result.stdout.strip()}
except Exception as e:
raise HTTPException(status_code=500, detail=str(e))
@router.post("/trends/run")
def trigger_trends_refresh():
"""手动刷新热点趋势数据"""
@@ -238,6 +255,7 @@ def get_modules_status():
today_str = date.today().isoformat()
log_based: dict = {
"scheduled_collect": {"name": "📡 内容采集", "log": LOGS_DIR / f"collector_{today_str}.log"},
"scheduled_refresh_search_cache": {"name": "🔍 搜索缓存", "log": LOGS_DIR / f"opencode_search_{today_str}.log"},
"scheduled_fetch_trends": {"name": "🔥 热点趋势", "log": LOGS_DIR / f"trends_{today_str}.log"},
"scheduled_generate": {"name": "🤖 内容创作", "log": LOGS_DIR / f"creator_{today_str}.log"},
"scheduled_optimize": {"name": "🔍 合规审查", "log": LOGS_DIR / f"optimizer_{today_str}.log"},
+33 -1
View File
@@ -62,6 +62,14 @@ class TaskScheduler:
max_instances=1,
coalesce=True
)
self.scheduler.add_job(
self._run_refresh_search_cache,
CronTrigger(hour=2, minute=30),
id='scheduled_refresh_search_cache',
replace_existing=True,
max_instances=1,
coalesce=True
)
self.scheduler.add_job(
self._run_metrics_sync,
CronTrigger(hour=6, minute=0),
@@ -72,7 +80,7 @@ class TaskScheduler:
)
self.scheduler.start()
self._started = True
logger.info("Scheduler started: 01:30 collect, 03:00 trends, 03:30 generate, 04:30 review, 05:00 optimize_sources, 06:00 metrics_sync")
logger.info("Scheduler started: 01:30 collect, 02:30 refresh_search, 03:00 trends, 03:30 generate, 04:30 review, 05:00 optimize_sources, 06:00 metrics_sync")
def shutdown(self):
if self.scheduler.running:
self.scheduler.shutdown()
@@ -97,6 +105,30 @@ class TaskScheduler:
except Exception as e:
logger.exception("[Scheduled] Trends refresh error: %s", e)
def _run_refresh_search_cache(self):
"""定时刷新搜索缓存(通过 opencode webfetch"""
try:
logger.info("[Scheduled] Refreshing search cache via opencode...")
import subprocess
result = subprocess.run(
[sys.executable, str(Path(__file__).parent.parent.parent.parent / "scripts" / "opencode_search.py"), "--refresh-cache"],
capture_output=True, text=True, timeout=600
)
for line in result.stdout.strip().split("\n"):
if line.strip():
logger.info("[SearchCache] %s", line.strip())
for line in result.stderr.strip().split("\n"):
if line.strip():
logger.warning("[SearchCache] %s", line.strip())
if result.returncode == 0:
logger.info("[Scheduled] Search cache refreshed")
else:
logger.warning("[Scheduled] Search cache refresh may have partial failures")
except subprocess.TimeoutExpired:
logger.warning("[Scheduled] Search cache refresh timed out")
except Exception as e:
logger.exception("[Scheduled] Search cache refresh error: %s", e)
def _run_generate(self):
try:
logger.info("[Scheduled] Starting content generation...")
+1
View File
@@ -285,6 +285,7 @@
this.runningModule = modId;
const endpoints = {
scheduled_collect: '/api/system/collect/run',
scheduled_refresh_search_cache: '/api/system/refresh-search-cache/run',
scheduled_fetch_trends: '/api/system/trends/run',
scheduled_generate: '/api/system/generate/run',
scheduled_optimize: '/api/system/review/run',
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env python3
"""
通过 opencode CLI 执行联网搜索
利用 opencode 的 webfetch 能力(当前 AI 环境可无障碍访问互联网)
用法:
python3 scripts/opencode_search.py --query "可持续生活 趋势 2026"
python3 scripts/opencode_search.py --refresh-cache # 刷新所有分类的缓存
"""
import argparse, datetime, json, logging, os, re, subprocess, sys, time
from pathlib import Path
from typing import Dict, List, Optional
PROJECT_ROOT = Path(__file__).parent.parent
sys.path.insert(0, str(PROJECT_ROOT))
LOGS_DIR = PROJECT_ROOT / "automation" / "logs"
TODAY = datetime.datetime.now().strftime("%Y-%m-%d")
logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(name)s - %(levelname)s - %(message)s')
logger = logging.getLogger(__name__)
SEARCH_CACHE_FILE = PROJECT_ROOT / "automation" / "data" / "search_cache.json"
def _run_opencode(prompt: str, timeout: int = 60) -> Optional[str]:
"""调用 opencode run 执行任务,返回文本输出"""
try:
result = subprocess.run(
["npx", "opencode", "run", prompt, "--format", "json"],
capture_output=True, text=True, timeout=timeout,
cwd=str(PROJECT_ROOT),
env={**os.environ, "OPENCODE_DISABLE_AUTOUPDATE": "1"}
)
if result.returncode != 0:
logger.warning(f"opencode run 返回非零: {result.stderr[:200]}")
return None
for line in result.stdout.strip().split("\n"):
try:
event = json.loads(line)
if event.get("type") == "error":
logger.warning(f"opencode 错误: {event}")
return None
except json.JSONDecodeError:
pass
lines = []
for line in result.stdout.strip().split("\n"):
try:
event = json.loads(line)
if event.get("type") == "text":
text = event.get("part", {}).get("text", "")
if text:
lines.append(text)
except json.JSONDecodeError:
pass
output = "\n".join(lines).strip()
return output if output else None
except subprocess.TimeoutExpired:
logger.warning(f"opencode run 超时 ({timeout}s)")
return None
except Exception as e:
logger.warning(f"opencode run 失败: {e}")
return None
def search_via_opencode(query: str, max_results: int = 5) -> List[Dict]:
"""通过 opencode 联网搜索"""
prompt = f"""用 webfetch 搜索:{query}
只输出 JSON 数组 [{{"title":"标题","url":"链接","content":"摘要"}}],最多 {max_results} 条,不要其他文字。"""
output = _run_opencode(prompt, timeout=90)
if not output:
return []
m = re.search(r'\[\s*\{.*\}\s*\]', output, re.DOTALL)
if not m:
logger.warning(f"未找到JSON数组: {output[:150]}")
return []
try:
results = json.loads(m.group())
if isinstance(results, list):
for r in results:
r["source"] = "opencode_webfetch"
logger.info(f"opencode 搜索 '{query[:20]}': {len(results)}")
return results[:max_results]
except Exception as e:
logger.warning(f"JSON解析失败: {e}")
return []
def refresh_cache():
"""刷新所有搜索分类的缓存"""
try:
with open(PROJECT_ROOT / "config" / "sources.yaml") as f:
import yaml
cfg = yaml.safe_load(f)
queries = [s["query"] for s in cfg["sustainability_sources"]["web_search"]]
except Exception:
logger.warning("无法读取 sources.yaml,使用默认查询")
queries = [
"以旧换新 二手交易 循环 2026",
"新能源车 绿色通勤 低碳 2026",
"干净饮食 有机食品 2026",
"零浪费 极简生活 可持续时尚 2026",
"绿色家电 一级能效 节能 2026",
"碳账户 碳普惠 个人碳减排 2026",
"环保科技 绿色产品 可持续材料 2026",
"AI工具 人工智能 效率提升 2026",
]
cache = {}
if SEARCH_CACHE_FILE.exists():
try:
cache = json.loads(SEARCH_CACHE_FILE.read_text(encoding="utf-8"))
except Exception:
pass
for i, q in enumerate(queries):
logger.info(f"[{i+1}/{len(queries)}] 搜索: {q}")
results = search_via_opencode(q, max_results=4)
if results:
cache[q] = results
else:
logger.warning(f" {q} 搜索无结果,保留旧缓存")
time.sleep(2)
SEARCH_CACHE_FILE.parent.mkdir(parents=True, exist_ok=True)
SEARCH_CACHE_FILE.write_text(json.dumps(cache, ensure_ascii=False, indent=2), encoding="utf-8")
logger.info(f"缓存已刷新: {sum(len(v) for v in cache.values())}")
def main():
parser = argparse.ArgumentParser(description="通过 opencode 联网搜索")
parser.add_argument("--query", help="搜索词")
parser.add_argument("--refresh-cache", action="store_true", help="刷新所有分类缓存")
parser.add_argument("--max-results", type=int, default=5)
args = parser.parse_args()
if args.refresh_cache:
refresh_cache()
return
if args.query:
results = search_via_opencode(args.query, args.max_results)
print(json.dumps(results, ensure_ascii=False, indent=2))
return
parser.print_help()
if __name__ == "__main__":
main()