Brandon Harley Dwiel
919-***-**** *******.*****@*****.*** http://www4.ncsu.edu/~bhdwiel/
EDUCATION
Doctor of Philosophy in Computer Engineering 2011-Present
North Carolina State University
Master of Science in Computer Engineering 2009-2011
North Carolina State University
Bachelor of Science in Computer Engineering 2004-2009
Southern Illinois University, Carbondale
Minor: Computer Science, Mathematics and Management
RESEARCH EXPERIENCE
2010 Present
Research Assistant, Dr. Rotenberg
Exploiting 3-D IC for Energy Efficient Heterogeneous Processors
Lead architecture researcher charged with designing heterogeneous processors that, when implemented using 3 -D IC
technology, delivers an improved performance/power spectrum over the 2-D version of the same design. Avenues of
research includes: designing and characterizing processor designs based on timing, power and area; fast thread migration
using through-silicon-vias; sharing structures (e.g., caches) among cores on different tiers; and exploiting on-chip
DRAM. The project includes yearly fabrications.
Comprehensively and Dynamically Reconfigurable Superscalar Processor
Lead researcher of a dynamically reconfigurable superscalar processor. Each configuratio n mimics the frequency,
performance and power of static designs while also dynamically adapting to changing workload behaviors.
FPGA Modeling of Diverse Superscalar Processors
Developed a tool to automatically construct, synthesize and simulate diverse superscalar processors on an FPGA.
Performance, resources and complexity are managed using various FPGA-specific techniques (e.g., clock decoupling).
Interfaces were developed in Verilog for connecting peripheral devices using a Wishbone bus.
2008 2009
Undergraduate Research Assistant, Dr. Zhang
Study on the Energy Efficiency of Just-in-time Compiler Optimizations
Researched the impact of dynamic compiler optimizations on the performance and energy consumption of mobile
devices. Energy dissipation was modeled using Wattch models. This research was presented at the Argonne National
Laboratory s Symposium for Undergraduates in Science, Engineering and Mathematics (2008).
PUBLICATIONS
N. K. Choudhary, S. Wadhavkar, J. Gandhi, T.A. Shah, H. Mayukh, Brandon H. Dwiel, S. Navada, H.H. Najaf-abadi,
and Eric Rotenberg. FabScalar: Composing Synthesizable RTL Designs of Arbitrary Cores within a Canonical
Superscalar Template, IEEE Micro, Special Issue: Micro's Top Picks from 2011 Computer Architecture Conferences
(MICRO TOP PICKS), Vol. 32, No. 3, May/June 2012.
Brandon H. Dwiel, N. K. Choudhary, and Eric Rotenberg. FPGA Modeling of Diverse Superscalar Processors, in
Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software, 2012.
N. K. Choudhary, S. Wadhavkar, J. Gandhi, T.A. Shah, H. Mayukh, Brandon H. Dwiel, S. Navada, H.H. Najaf-abadi,
and Eric Rotenberg. FabScalar: Composing Synthesizable RTL Designs of Arbitrary Cores within a Canonical
Superscalar Template, in Proceedings of the 38th International Symposium on Computer Architecture (ISCA), 2011 .
SKILLS
Tools: Xilinx ISE, Cadence (NC-Verilog, Virtuoso, Encounter), Synopsys (Design Compiler), SPICE, Mentor (ModelSim,
Questa), Universal Verification Methodology/Open Verification Methodology, SimpleScalar
Programming Languages: Verilog, SystemVerilog, VHDL, C, C++, System C, Java, OpenMP, MPI, CUDA, OpenCL,
Perl, Python, Linux shell and Linux system programming
RELATED COURSE PROJECTS
Implementation of a novel, copy-free RMT recovery scheme (C++ and Verilog)
Implementation of an image rotation function (CUDA and OpenCL)
Schematic to final layout of an 8x8 CAM using 45 nm technology (optimized for )
Implementation of an out-of-order superscalar pipeline simulator based on Tomasulo s algorithm (C++)
Implementation and physical design of a variable block size motion estimator (VBSME) for low -power H.264 video
compression (Verilog)
Frontend and backend compiler implementation for the ICE9 academic ISA (Flex, Bison, C)
Implementation of a multi-core cache simulator with the MESI cache coherence protocol (C)
Parallelization of various programs using OpenMP, MPI, CUDA and the Cell processor ( C).
RELATED GRADUATE COURSE WORK
Computer Architecture: Advanced Microarchitecture, Computer Design and Technology, Architecture of Parallel
Computers, Multi-Core/Many-Core GPU Architecture and Programming
VLSI: Electronic System Level and Physical Design, Digital Electronics, Digital ASIC Design, VLSI Systems Design
Programming: Programming Parallel Systems, Embedded Systems with Linux and Android, Compiler Construction
HONORS AND ACHIEVEMENTS
Presenter, Argonne Symposium for Undergraduates in Science, Eng ineering and Mathematics 2008
Sigma Alpha Lambda, National Leadership and Honors Fall 2006 - Present