Roger Kim

Roger Kim Python과 Verilog를 활용한 실전 중심의 AI NPU 설계를 공유하며, 첨단 반도체 시스템 설계를 더 쉽고 널리 배울 수 있도록 돕고 있습니다. RISC-V CPU, AI NPU, FPGA 기반 Edge AI 시스템을 연구·개발하고 있습니다.

그동안 개발해 온 RISC-V CPU 설계 교육 시리즈의 Volume 3를 완성하여 소개드립니다.이번 과정에서는 Volume 2의 TinyRV32I에서 한 단계 더 나아가 RV32I 40개 명령어를 사용하는 Full...
27/08/2026

그동안 개발해 온 RISC-V CPU 설계 교육 시리즈의 Volume 3를 완성하여 소개드립니다.

이번 과정에서는 Volume 2의 TinyRV32I에서 한 단계 더 나아가 RV32I 40개 명령어를 사용하는 Full CPU를 직접 설계합니다.

5단 Pipeline을 기반으로 Forwarding과 Load-Use Interlock을 구현하고, CSR, Exception, Interrupt, Trap까지 통합하여 실제 프로그램을 실행할 수 있는 CPU를 완성합니다.

CPU를 설계하는 것만큼 “내가 만든 CPU가 정말 올바르게 동작하는가?”를 검증하는 과정도 중요하게 다루었습니다.

Testbench와 Simulation, 실제 프로그램 실행, Spike Reference Model 비교를 거쳐 RISC-V Architecture Compatibility Test를 적용했습니다. 총 48개 시험 중 47 PASS, 1 Waiver의 결과와 기술적 근거도 RISC-V 공식 GitHub에 공개되어 있습니다.

마지막에는 검증한 CPU를 Arty S7-25 FPGA에 직접 구현하여 실제 프로그램이 동작하는 것까지 확인합니다.

40 Instructions → 5-Stage Pipeline → Hazard Control → CSR → Exception & Interrupt → Trap → Program Ex*****on → Spike Verification → ACT → FPGA

CPU를 이론으로만 설명하기보다 명령어에서 CPU 구조를 도출하고, Verilog RTL로 직접 설계하고, 검증한 뒤 실제 FPGA에서 프로그램을 실행하는 전체 과정을 연결해서 보여주는 것이 이번 강의에서 가장 중요하게 생각한 부분입니다.

🎓 FPGA 구현을 위한 RISC-V CPU 설계 ③
인프런 전체 강의 보기: https://inf.run/G1auY

✅ EdgeChipLab_CPU_v1 · RV32I ACT 검증 결과
RISC-V GitHub에 공개된 Compliance Report:
https://github.com/riscv/riscv-arch-test/issues/1600

반도체·CPU·FPGA 설계에 관심 있는 학생과 엔지니어분들께 조금이나마 도움이 되었으면 합니다.

#반도체 #컴퓨터구조

PPACT Studio — Public PreviewI am pleased to share the Public Preview of PPACT Studio, a system-level architecture explo...
09/08/2026

PPACT Studio — Public Preview

I am pleased to share the Public Preview of PPACT Studio, a system-level architecture exploration and decision-support platform for AI systems.

PPACT stands for Performance, Power, Area, Cost, and Traffic — five system-level dimensions used to examine architecture trade-offs.

PPACT Studio goes beyond simply analyzing a design. It examines the interactions among CPUs, AI accelerators, memory, data movement, power, silicon area, and cost to identify bottlenecks and trade-offs. It then provides engineering reasons for the results and recommends meaningful architecture changes for further evaluation.

The results are not presented as numbers alone. PPACT Studio generates multiple engineering views, including System Flow diagrams, Spider Maps, Balance views, measured-result comparisons, bottleneck analysis, and engineering conclusions, so that users can see not only what changed, but also why it changed.

Rather than reducing AI system performance to a single number such as TOPS, PPACT Studio distinguishes among single-job latency, sustainable pipeline capacity, and delivered throughput.

The objective is not to replace detailed simulation or EDA tools. PPACT Studio is intended to support an earlier stage of architecture exploration, helping engineers compare alternatives, understand system-level interactions, and identify which design directions deserve deeper investigation.

PPACT Studio is currently in Public Preview. At this stage, I would especially value technical feedback from engineers, researchers, educators, and anyone interested in AI hardware and system architecture.

If you have a question, suggestion, technical concern, or a different interpretation of the results, please leave a comment on the YouTube video. I will give feedback posted there priority, review it carefully, and respond as promptly as possible.

Comments on modeling assumptions, engineering interpretations, visualizations, or areas that could be improved are particularly welcome. Critical feedback will help strengthen the platform as it develops toward version one point zero.

🎥 Watch PPACT Studio
— https://www.youtube.com/watch?v=0fNdJgugudw

PPACT Studio — Compare. Explain. Decide.

28/07/2026

안녕하세요. 반도체 설계 교육 커리큘럼을 시작하며, 지금까지 만든 첫 다섯 편을 먼저 공유합니다.

삼성전자에서 30여 년간 실무를 경험하고, 현재 대학에서 학생들을 가르치며 늘 마음속에 두었던 목표가 있었습니다. 트랜지스터부터 CPU, NPU, AI SoC, 그리고 소형 LLM에 이르기까지, 단편적인 이론에 그치지 않고 직접 설계하여 실제 FPGA에서 동작시켜보는 완결된 커리큘럼을 만드는 것입니다.

이 과정의 가장 중요한 기준은 단순히 "동작한다"를 넘어, 소프트웨어 모델과 하드웨어 결과가 1비트의 오차도 없이 "정확히 일치한다(Bit-True)"는 것을 증명하는 데 있습니다. GitHub에 공개된 소스코드는 검증되고 온전히 동작하는 코드로써 누구나 다운로드 받아서 사용할 수 있습니다.

■ 현재 공개된 다섯 편의 과정

▪ AI 반도체 설계를 위한 필수 이론 — 뉴런에서 LLM까지
하드웨어 설계에 실질적으로 필요한 AI 이론만을 압축하여 담았습니다.
https://inf.run/VajoB

▪ FPGA 구현을 위한 머신러닝 파이썬 실습
소프트웨어 모델과 실제 칩 사이의 간극을 메우는 결정론적 모델링 과정입니다.
https://inf.run/wCWD1

▪ FPGA로 만드는 실전 하드웨어 — 디지털 시계부터 실시간 카메라까지
7세그먼트, 초음파 계측, 카메라 영상처리 파이프라인 등을 직접 구현합니다.
https://inf.run/ujcbB

▪ FPGA로 만드는 AI 가속기 — 이미지 처리부터 하드웨어 검증까지
CNN 학습부터 INT8 양자화, 그리고 FPGA Bit-True 검증의 전 과정을 다룹니다.
https://inf.run/nTLKT

▪ FPGA로 만드는 MNIST NPU — RTL부터 Bit-True 검증까지
파이썬 참조 모델과 Verilog RTL의 100% 일치를 증명하는 NPU 설계의 집약체입니다.
https://inf.run/efRnt

■ 앞으로 준비 중인 심화 과정

▪ RISC-V CPU 설계 (공식 컴플라이언스 테스트 통과 과정)
▪ NPU 심화 (CIFAR-100, 쿼드코어 Systolic Array 구조)
▪ AI SoC 통합 (CPU + NPU + 주변장치 통합 시스템)
▪ 실시간 비전 시스템 및 소형 LLM 트랜스포머 가속 엔진

모든 과정은 학습의 문턱을 낮추기 위해 보급형 보드(Arty S7-25)와 무료 툴(Vivado, 파이썬)만으로 재현할 수 있도록 구성하였으며, 전체 소스 코드는 깃허브를 통해 열어두었습니다.

반도체 설계를 근본부터 다지고자 하는 학생분들과 현업에서 팀 단위의 교육을 고민하시는 분들께 작은 보탬이 되기를 바랍니다. 남은 과정들도 묵묵히 준비하여 이어가겠습니다. 감사합니다!

▪ 공식 실습 소스코드 (GitHub): https://github.com/estlit/SemiconductorSchool-Labs
▪ 아키텍처 해설 (YouTube):

#반도체설계 #시스템반도체 #차세대반도체

신경망을 FPGA에 올린다 — 말은 익숙하지만, 실제로 해낸 사람은 드뭅니다.MNIST 손글씨 숫자를 인식하는 신경망 가속기(NPU)를, 스펙 읽기부터 파이썬 골든 모델, 파이썬 학습하여 가중치 산출, Verilog...
25/07/2026

신경망을 FPGA에 올린다 — 말은 익숙하지만, 실제로 해낸 사람은 드뭅니다.

MNIST 손글씨 숫자를 인식하는 신경망 가속기(NPU)를, 스펙 읽기부터 파이썬 골든 모델, 파이썬 학습하여 가중치 산출, Verilog RTL 설계, 그리고 FPGA에서 "파이썬 정답지와 한 비트도 다르지 않음(Bit-True)"을 증명하는 검증까지 — 처음부터 끝까지 직접 만듭니다.

딥러닝 이론은 몰라도 됩니다. Verilog 기본기만 있으면 따라올 수 있도록 설계했습니다. 캡스톤·졸업작품·경진대회에 그대로 낼 수 있는 완성된 하드웨어 프로젝트가 손에 남습니다.

보급형 FPGA(Arty S7-25) 한 장과 무료 툴(Vivado)만으로, 모든 소스는 GitHub 공개.

🎓 FPGA로 만드는 MNIST NPU — RTL부터 Bit-True 검증까지
🎁 오픈 기념 할인 중
▶ https://inf.run/efRnt

#반도체설계 #캡스톤 #임베디드

"이론은 아는데 직접 설계해본 적이 없어서 자소서에 어필할 내용이 없다."반도체 설계 취업을 준비하는 학생들에게 가장 많이 듣는 말입니다.삼성전자에서 30년, 지금은 대학에서 반도체 설계를 가르치며 준비한 실습 강의...
23/07/2026

"이론은 아는데 직접 설계해본 적이 없어서 자소서에 어필할 내용이 없다."

반도체 설계 취업을 준비하는 학생들에게 가장 많이 듣는 말입니다.
삼성전자에서 30년, 지금은 대학에서 반도체 설계를 가르치며 준비한 실습 강의를 인프런에 열었습니다. 보드 한 장과 무료 툴로, 면접에서 "제가 직접 설계하고 검증했습니다"라고 말할 수 있는 결과물을 만드는 과정입니다.

▶ FPGA로 만드는 AI 가속기 (CNN 학습 → 양자화 → Bit-True 검증)
https://inf.run/nTLKT
▶ FPGA 실전 하드웨어 (시계·초음파·OLED·실시간 카메라 4개 프로젝트)
https://inf.run/ujcbB
▶ AI 반도체 설계를 위한 필수 이론 (뉴런에서 LLM까지)
https://inf.run/VajoB

모든 소스는 GitHub에 공개되어 있고, 보급형 FPGA(Arty S7-25)와 무료 Vivado만으로 완주할 수 있습니다.

#반도체취업 #팹리스 #전자공학

Building an Open FPGA & AI SoC Education PlatformOver the past several years, I have been developing Semiconductor Schoo...
10/07/2026

Building an Open FPGA & AI SoC Education Platform

Over the past several years, I have been developing Semiconductor School, an open educational platform covering the complete semiconductor design flow—from Digital Logic to AI SoC implementation.
Unlike conventional lecture-based courses, every project follows the complete hardware development flow:

{Python Modeling → Verilog RTL Design → Testbench Verification → FPGA Implementation}

All Verilog HDL source code, Python reference models, and testbenches are freely available on GitHub, allowing learners to reproduce every project on their own FPGA boards.

The curriculum is organized into 10 progressive levels.
✅ Currently Available
[Level 1] Digital Design Fundamentals
- What is an FPGA?
- Digital Design Flow
- Verilog HDL
- Introduction to AI Semiconductors
[Level 2] FPGA Hands-on Labs
- 7-Segment Digital Clock
- Smart Ultrasonic Distance Meter
- OLED Display
- Real-Time Camera Interface
- Level 5 – AI Hardware
- NPU Systolic Processing Element

✅ Currently in Production (FPGA Implementation Completed)
The following projects have already been designed, implemented, and verified on FPGA. Lecture videos are currently being produced and will be released sequentially.
- Advanced Verilog HDL
- Finite State Machine (FSM)
- UART / SPI / I²C Interfaces
- VGA / HDMI Display Controller
- FPGA Image Processing
- RISC-V RV32I 5-Stage Pipeline CPU
- RISC-V Compliance Verification
- AXI Bus Architecture
- Embedded SoC Design
- CNN-based AI NPU
- Transformer Hardware Accelerator
- Mini LLM Hardware
- FPGA-based AI SoC Integration

My goal is to provide a complete educational pathway from Digital Logic Design to AI SoC Design, enabling students, educators, and engineers to learn by building real hardware.
Although the lectures are presented in Korean, English subtitles are available via YouTube auto-translation, and all source code is openly available on GitHub.
I hope these resources will be useful to the global FPGA, RISC-V, and AI hardware communities.
Feedback and suggestions are always welcome.

📺 YouTube
https://youtube.com/
💻 GitHub
https://github.com/estlit

[AI 반도체의 비밀, 엑셀로 직접 돌려봅니다]AI 연산을 획기적으로 빠르게 만드는 NPU의 심장, 시스톨릭 어레이(Systolic Array).어렵고 복잡한 수식과 하드웨어 구조를 '자동 세차장' 비유와 '엑셀 실...
28/06/2026

[AI 반도체의 비밀, 엑셀로 직접 돌려봅니다]

AI 연산을 획기적으로 빠르게 만드는 NPU의 심장, 시스톨릭 어레이(Systolic Array).
어렵고 복잡한 수식과 하드웨어 구조를 '자동 세차장' 비유와 '엑셀 실습'으로 완벽하게 분해했습니다!

기존 CPU의 일반적인 메모리 접근 방식과 NPU의 대량 병렬 처리 방식은 과연 얼마나 차이가 날까요? 3x3 행렬 기준으로 무려 약 3.9배(27 step vs 7 cycle)의 속도 차이가 발생합니다. 행렬이 커질수록 그 격차는 상상을 초월하죠.

Semiconductor School이 준비한 이번 영상에서는, 각 PE(Processing Element) 블록 사이로 데이터가 어떻게 흘러가며 연산되는지, 엑셀 실습 파일을 통해 여러분이 직접 눈으로 확인하고 검증할 수 있습니다.
유투브 영상을 올립니다.

💡 지금 바로 영상 확인하고 실습 파일도 다운로드해 보세요!
https://youtu.be/sS157Ydtmfg?si=zSG7_A8kGr9JMhf7

📂 실습용 엑셀 파일 다운로드
https://github.com/estlit/AI_NPU_System_Design_v1/blob/main/MNIST_System_Design.zip

📖 더 깊이 공부하고 싶다면 (Amazon Best Seller #2 in Microprocessor Design):
https://www.amazon.com/dp/B0GMPSND15

#반도체 #하드웨어설계 #딥러닝하드웨어

Semiconductor School | EdgeChipLab

안녕하세요! 차세대 칩 설계의 기준, Semiconductor School입니다.AI 반도체, NPU(Neural Pr...

오랜 기간 준비해 온 프로젝트가 조금씩 모습을 갖춰가고 있습니다. 반도체 소자부터 RISC-V CPU, AI NPU, AI SoC, Transformer, 그리고 Mini LLM에 이르기까지 전체 시스템 스택을 직접...
27/06/2026

오랜 기간 준비해 온 프로젝트가 조금씩 모습을 갖춰가고 있습니다.

반도체 소자부터 RISC-V CPU, AI NPU, AI SoC, Transformer, 그리고 Mini LLM에 이르기까지 전체 시스템 스택을 직접 설계하고 구현하는 교육 플랫폼 EdgeChipLab을 준비하고 있습니다. 준비하고 있는 내용으로 영상을 만들었습니다.

하나씩 검증된 자산들이 쌓여가고 있고 플랫폼의 윤곽도 점차 분명해지고 있습니다. 2026년 12월 공식 오픈을 목표로 준비하고 있습니다.
유투브 영상으로 소개드립니다.
실전형 반도체 설계 교육의 새로운 기준을 만들어 보고자 합니다!

I’m happy to share a meaningful milestone from our open hardware education project at EdgeChipLab.Our custom pure-Verilo...
30/05/2026

I’m happy to share a meaningful milestone from our open hardware education project at EdgeChipLab.

Our custom pure-Verilog RV32I CPU core, EdgeChipLab_CPU_v1, has successfully passed the RISC-V RV32I Architectural Compatibility Tests (ACT v3.0).

Highlights:
• Pure RTL educational CPU core
• Cycle-by-cycle bit-true regression against Spike
• 48 architectural tests
• 47 passed + 1 waived (misaligned jump corner case with detailed public analysis)
• Public acknowledgement from the official RISC-V Arch-Test maintainer

GitHub Issue: https://github.com/riscv/riscv-arch-test/issues/1600
Compliance Repository: https://github.com/estlit/EdgeChipLab_RV32I_Compliance_Report

This milestone is especially meaningful because it directly connects to our broader educational roadmap:

RTL Design → RISC-V CPU → Architectural Verification → NPU Integration → FPGA-based Edge AI System

At EdgeChipLab, our goal is to make advanced semiconductor system design transparent and teachable through complete hands-on implementation.

The newly verified EdgeChipLab_CPU_v1 is planned to be integrated with our custom RTL-based AI NPU into a complete memory-mapped SoC platform on FPGA.

This next-stage integration is not simply connecting CPU and accelerator blocks—it requires a deep understanding of how an NPU processes data internally, how memory interfaces are organized, and how bit-true hardware verification is maintained across the entire system.

For students and engineers interested in following this development, I recommend studying the NPU architecture and verification flow presented in my Amazon book:

“AI NPU System Design with Python and Verilog”
by Roger Kim (Hyo Seob Kim)

Amazon: https://www.amazon.com/dp/B0GLQVJWMK

The book explains how a custom RTL-based neural processing unit can be designed from scratch, verified bit-by-bit using Python and Verilog, and implemented on FPGA.

This foundation becomes especially valuable as we move toward:

Verified RISC-V CPU + Custom NPU + Memory-Mapped SoC + FPGA Implementation

—from architecture fundamentals to a complete edge AI semiconductor system.

My hope is that students and engineers can learn not only individual modules, but the full engineering path of building, integrating, and verifying practical semiconductor systems from scratch.

Dear RISC-V Arch-Test Team, We are pleased to report that our custom pure RV32I RTL core, EdgeChipLab_CPU_v1, has successfully passed 100% of the Base Integer Architectural Certification Tests. We ...

🌍 전 세계 엔지니어들과 함께하는 AI NPU 설계의 여정!아마존 베스트셀러 2위에 오른 저의 저서 'AI NPU System Design with Python and Verilog'가 이제 미국, 스페인, 브라질을...
21/04/2026

🌍 전 세계 엔지니어들과 함께하는 AI NPU 설계의 여정!

아마존 베스트셀러 2위에 오른 저의 저서 'AI NPU System Design with Python and Verilog'가 이제 미국, 스페인, 브라질을 넘어 전 세계 곳곳의 엔지니어와 연구자들에게 읽히고 있습니다. 아래 사진은 이번달 책이 판매된 전세계 나라들입니다.

단순한 이론에 머무르지 않고, '실제로 동작하는 실리콘(Working Silicon)'을 구현하기 위한 저의 설계 철학이 언어의 장벽을 넘어 글로벌 시장에서도 깊은 공감을 얻고 있어 매우 뜻깊게 생각합니다.
각국의 독자들이 보내주시는 뜨거운 성원에 힘입어, 앞으로도 대한민국 반도체 설계 기술의 저력을 세계 시장에 널리 알리고 차세대 AI 하드웨어 생태계를 구축하는 데 정진하겠습니다. 지금 바로 아마존과 깃허브에서 확인해 보세요!

※ 책을 구매하지 않으셔도 파이썬/Verilog Source code 전체를 무료로 다운로드 받을수 있습니다.

🛒 Amazon (Global): https://www.amazon.com/dp/B0GLQVJWMK
💻 GitHub (Free Download Python/Verilog Full Sources): https://github.com/estlit

Address

서울특별시 동작구 상도로 369 숭실대학교
Seoul
06978

Alerts

Be the first to know and let us send you an email when Roger Kim posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Contact The Business

Send a message to Roger Kim:

Shortcuts

Share