<?xml version='1.0' encoding='utf-8'?>
<rss version="2.0">
  <channel>
    <title>Wukai’s Lab</title>
    <link>https://kwu130.github.io/</link>
    <description>C++、CUDA 与工程实践笔记</description>
    <language>zh-CN</language>
    <item>
      <title>从函数拦截到调用计时：用 LD_PRELOAD 写一个 API Profiler</title>
      <link>https://kwu130.github.io/posts/ld-preload-profiler/</link>
      <guid isPermaLink="true">https://kwu130.github.io/posts/ld-preload-profiler/</guid>
      <description>用一个 C++ 动态库实验，理解 LD_PRELOAD 如何拦截 API、函数别名如何保留原始入口，以及怎样测量和解释调用耗时。</description>
      <pubDate>Thu, 01 Oct 2026 00:00:00 +0000</pubDate>
      <category>C/C++</category>
      <category>调试</category>
      <category>性能优化</category>
    </item>
    <item>
      <title>NVLS：NCCL 如何借助 NVSwitch 做网络内归约</title>
      <link>https://kwu130.github.io/posts/nvls/</link>
      <guid isPermaLink="true">https://kwu130.github.io/posts/nvls/</guid>
      <description>从 AllReduce 的数据路径、NCCL 的算法选择到 CUDA multicast 映射，梳理 NVLink SHARP 的工作边界、验证方法与常见误区。</description>
      <pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate>
      <category>CUDA</category>
      <category>集合通信</category>
      <category>性能优化</category>
    </item>
    <item>
      <title>CUDA Shared Memory Bank Conflict：从地址映射到 Padding 与 XOR Swizzling</title>
      <link>https://kwu130.github.io/posts/cuda-shared-memory-bank-conflicts/</link>
      <guid isPermaLink="true">https://kwu130.github.io/posts/cuda-shared-memory-bank-conflicts/</guid>
      <description>从 shared memory 的 bank 映射和冲突度出发，用矩阵转置解释 32-way conflict，并比较 Padding、XOR Swizzling 与 Nsight Compute 验证方法。</description>
      <pubDate>Sun, 30 Aug 2026 00:00:00 +0000</pubDate>
      <category>CUDA</category>
      <category>性能优化</category>
    </item>
    <item>
      <title>用 Valgrind 定位 C/C++ 内存问题：泄漏、越界与未初始化读取</title>
      <link>https://kwu130.github.io/posts/valgrind-memcheck/</link>
      <guid isPermaLink="true">https://kwu130.github.io/posts/valgrind-memcheck/</guid>
      <description>从可复现示例出发，读懂 Memcheck 的泄漏分类、非法访问和未初始化值报告，并把检查接入日常开发流程。</description>
      <pubDate>Sat, 27 Jul 2024 00:00:00 +0000</pubDate>
      <category>C/C++</category>
      <category>调试</category>
    </item>
    <item>
      <title>CUDA 设备代码中的 long double：为什么不能把它当作扩展精度</title>
      <link>https://kwu130.github.io/posts/cuda-device-long-double/</link>
      <guid isPermaLink="true">https://kwu130.github.io/posts/cuda-device-long-double/</guid>
      <description>从 CUDA 的语言支持边界、主机 ABI 和对象表示出发，重新审视 long double 在设备代码中的警告与实验结果。</description>
      <pubDate>Sat, 27 Jul 2024 00:00:00 +0000</pubDate>
      <category>C/C++</category>
      <category>CUDA</category>
    </item>
    <item>
      <title>VS Code 开发环境配置：C/C++、Python 与远程开发</title>
      <link>https://kwu130.github.io/posts/vscode-development-setup/</link>
      <guid isPermaLink="true">https://kwu130.github.io/posts/vscode-development-setup/</guid>
      <description>用一组职责清晰的扩展和少量可迁移设置，搭建适合 C/C++、Python、CMake 与 Remote SSH 的开发环境。</description>
      <pubDate>Mon, 01 Jul 2024 00:00:00 +0000</pubDate>
      <category>开发工具</category>
      <category>VS Code</category>
    </item>
  </channel>
</rss>