To investigate the cellular diversity across human cortex, a low-bias approach to profile cell-type diversity was sought, constrained by the challenge of working with precious and limited tissue sources. Individual layers of cortex were dissected from tissues covering the middle temporal gyrus (MTG), anterior cingulate gyrus (CgGr), primary visual cortex (V1C), primary motor cortex (M1C), primary somatosensory cortex (S1C) and primary auditory cortex (A1C) derived from human brain, and nuclei were dissociated and sorted using the neuronal marker NeuN. Nuclei were sampled from postmortem and neurosurgical (MTG only) donor brains, and expression was profiled with SMART-Seq v4 or 10x v3 RNA-sequencing.
This database of cell types includes experimental data derived from adult human brain. Human brain tissue samples from either postmortem or neurosurgical origin were made available through the generosity of tissue donors. Clinical summaries and donor characteristics are provided in this document as well as a description of the criteria for acceptance of use in this study.
To prepare and archive tissues from suitable cases, whole postmortem brain specimens were bisected through the midline, and individual hemispheres were embedded in alginate for slabbing. Coronal brain slabs were cut at 0.5-1cm intervals through each hemisphere and the slabs were then frozen in a bath of dry ice and isopentane, vacuum sealed in freezer bags to prevent frost damage, and stored at -80°C until use. Regions of interest were subsequently removed from tissue slabs, sectioned on a vibratome and processed for nuclei isolation.
Neurosurgical donor tissue (MTG only) was received from patients undergoing surgery for epilepsy or brain tumors. The tissue blocks received were distal, apparently normal cortical tissue removed to access underlying pathological brain tissues. Tissue was transported in chilled ACSF, sectioned at 350µm and stored at -80°C until they were processed for nuclei isolation.
Single nuclei were captured by gating on DAPI-positive events, excluding debris and doublets, and then gating on NeuN signal, which allowed for the isolation of either NeuN-positive (neuronal) or NeuN-negative (non-neuronal) events.
PROTOCOL Human Tissue Sectioning and Dissection for Nuclear Isolation
PROTOCOL Isolation of Nuclei from Adult Human Brain Tissue
PROTOCOL Isolation of Nuclei from Adult Brain Tissue from 10x Genomics Platform
SMART-Seq v4 Ultra Low Input RNA Kit for Sequencing (Takara #634894) was used per the manufacturer’s instructions for cDNA synthesis of single-cell RNA and subsequent amplification. Sequencing libraries were prepared using the NexteraXT DNA Library Preparation kit (Illumina FC-131-1096) with NexteraXT Index Kit V2 Set A, B, C, or D (FC-131-2001, 2002, 2003, or 2004) or custom 8-base or 10-base Unique Design index primers designed and manufactured by IDT (Integrated DNA Technologies). NexteraXT DNA Library prep was done at either 0.5x volume manually or 0.4x or 0.2x volume on the Mantis instrument (Formulatrix). Pooled sequencing libraries were sent to an outside vendor for sequencing on an Illumina HiSeq 2500 instrument. All of the library pools were run using Illumina High Output V4 chemistry. RNA sequencing services were provided by Covance Genomics Laboratory, Seattle subsidiary of LabCorp Group of Holdings, and The Broad Institute Genome Sequencing Platform.
PROTOCOL Nextera XT at 0.2X on the Mantis
PROTOCOL SMART-Seq v4 (1x) amplification
PROTOCOL SMART-Seq v4 (0.5x) amplification
Raw read (fastq) files were aligned to the GRCh28 human genome sequence (Genome Reference Consortium, 2011) with the RefSeq transcriptome version GRCh28.p2 (current as of 4/13/2015) and updated by removing duplicate Entrez gene entries from the gtf reference file for STAR processing. For alignment, Illumina sequencing adapters were clipped from the reads using the fastqMCF program. After clipping, the paired-end reads were mapped using Spliced Transcripts Alignment to a Reference (STAR) using default settings. Reads that did not map to the genome were then aligned to synthetic construct (i.e. ERCC) sequences and the E. coli genome (version ASM584v2). Quantification was performed using summerizeOverlaps from the R package GenomicAlignments. Expression levels were calculated as counts per million (CPM) of exonic plus intronic reads.
Single nucleus suspensions were frozen in a solution of 1X PBS, 1% BSA, 10% DMSO, and 0.5% RNAsin Plus RNase inhibitor (Promega, N2611) and stored at -80°C. At the time of use, frozen nuclei were thawed at 37°C and processed for loading on the 10x Chromium instrument as described (dx.doi.org/10.17504/protocols.io.nx3dfqn). Samples were processed using the 10x Chromium Single Cell 3’ Reagent Kit v3. 10x chip loading and sample processing was done according to the manufacturer’s protocol. Gene expression was quantified using the default 10x Cell Ranger v3 pipeline except substituting the curated genome annotation used for SMART-seq v4 quantification. Introns were annotated as “mRNA,” and intronic reads were included in expression quantification.
Nuclei were included in the clustering analysis if they passed all QC criteria.
SMART-seq v4 criteria:
10x v3 criteria:
Nuclei passing QC criteria were grouped into transcriptomic cell types using an iterative clustering procedure previously reported in (Tasic et al. 2018; Hodge, Bakken et al., 2019). Briefly, intronic and exonic read counts were summed, and log2-transformed expression was centered and scaled across nuclei. X- and Y-chromosomes and mitochondrial genes were excluded to avoid nuclei clustering based on sex or nuclei quality. Differentially expressed genes were selected, principal components analysis (PCA) reduced dimensionality, and a nearest neighbor graph was built using up to 20 principal components. Clusters were identified with Louvain community detection (or Ward's hierarchical clustering if N < 3000 nuclei), and pairs of clusters were merged if either cluster lacked marker genes. Clustering was applied iteratively to each sub-cluster until clusters could not be further split.
Cluster robustness was assessed by repeating iterative clustering 100 times for random subsets of 80% of nuclei. A co-clustering matrix was generated that represented the proportion of clustering iterations that each pair of nuclei were assigned to the same cluster. We defined consensus clusters by iteratively splitting the co-clustering matrix as described (Tasic et al. 2018; Hodge, Bakken et al., 2019).
Clusters were curated based on outlier values of the initial QC values or cell class marker expression (GAD1, SLC17A7, SNAP25). Clusters were identified as donor-specific if they included fewer nuclei sampled from donors than expected by chance. To confirm exclusion, clusters automatically flagged as outliers or donor-specific were manually inspected for expression of broad cell class marker genes, mitochondrial genes related to quality, and known activity-dependent genes.
The clustering pipeline is implemented in the R package “scrattch.hicat”, and the clustering method is provided by the “run_consensus_clust” function.
CODE Hierarchical, iterative clustering for analysis of transcriptomics data in R
Data generation was supported by multiple awards, including Brain Initiative Cell Census Network (BICCN) award U01MH114812 from the National Institute of Mental Health and the National Institute of Neurological Disorders and Stroke, and by the Allen Institute for Brain Science.
Hodge, R.D., Bakken, T.E., et al. (2019). "Conserved cell types with divergent features in human versus mouse cortex." Nature 573:61-68. PMID DOI
Tasic, B., et al. (2018). "Shared and distinct transcriptomic cell types across neocortical areas." Nature 563(7729): 72-78. doi: 10.1038/s41586-018-0654-5. Epub 2018 Oct 31. PMID PMCID DOI
This third section "Putting it all together" is a key to link research focus to related scientific tool to primary paper to taxonomy used.
This table links the research topic (species, brain area of interest, and technique) to the related tool, the related dataset, the related paper, and the related taxonomy.
Discover all pages in this series:
This second section "How to..." lists more technical resources such as tutorials, protocols, use cases, and more.
These are Allen Institute resources to explain how to use our datasets, replicate our experiments with protocols, and links to tables with important metadata and features definitions. These sections are designed for scientists who are interested in more technical support in their cell types research.
Identifying and naming brain cells has been an integral part of neuroscience for a century, including Allen Institute cell typing efforts. This means many cell types have multiple names, and tracking this information requires standards.
1. Understand the Allen Institute framework for developing nomenclatures called the Common Cell Type Nomenclature, or CCN.
Related publication (Miller et al 2020)
2. Learn about the schema the Allen Institute has developed for defining cell type taxonomy components, such as nomenclature, annotations, metadata, and the underlying transcriptomic data, or use this format for your own data.
Allen Institute Taxonomies in AIT Format
3. Compare how cell type names have changed as the Allen Institute collects more data from across the brain, and see how brain region, age, gene counts, and other cell features relate to cell types.
Annotation Comparison Explorer
Adopting the Patch-seq technique in a lab can be daunting, but these protocols and resources can help. The Allen Institute has adopted and extended multiple experimental and computational protocols to make Patch-Seq is a useful method for understand the structure and function of brain defined cell types.
These resources below are used in multiple tools and data sets on Allen Brain Map:
Related tool: Allen Cell Types database
Related tool: Mouse PatchSeq VISp viewer
Related tool: Cell Type Knowledge Explorer
1. This GitHub provides a starting point for labs interested in using the Patch-seq technique or refining their existing technique. Specifically, this resource consists of three components: (1) a step-by-step optimized Patch-seq protocol (2) the Multichannel Igor Electrophysiology Suite (MIES) software package and (3) an R library that uses a modified workflow.
Related publication (Lee et al 2021)
2. Overview of the various steps and tools to generate data across species and brain regions in Patch-Seq.
3. From our electrophysiology team, this is detailed protocol to obtain electrophysiological recordings and cellular contents from neurons in postnatal mouse and/or human brain slices.
Related publication (Lee et al 2021)
4. IPFX is a Python package for computing intrinsic cell features from electrophysiology data. That can perform cell data quality control (e.g. resting potential stability), detect action potentials and their features (e.g. threshold time and voltage), calculate features of spike trains (e.g., adaptation index), and calculate stimulus-specific cell features.
Automatic computing for electrophysiological cell features
5. Here is your one-stop shop for the Allen Institute’s free, open-source neuron reconstruction software, protocols, and analysis scripts for generating and analyzing image-based, quantitative, 3D morphologies for your own research.
Protocol, Software, Analysis: Reconstruct neuron morphology
6. Protocol to generate full-length cDNA from single cells, or nuclei, using Takara SMARTer V4.
Takara SMARTer V4 (Protocols.io protocol)
7. The electrophysiology and morphology feature definitions are used for all our Patch-Seq datasets. This is from “NIHMS1691616-supplement-Supplementary_Figures” in Gouwens, Sorensen, Berg, et al. 2019, starts at page 73 for electrophysiology and page 75 for morphology.
Electrophsyiology & Morphology Features Definition
Related publication (Gouwens et al 2019)
We define cell types based on which genes are turned on and which genes are turned off in a cell. These resources detail how we do this robustly for millions of cells.
1. This page includes protocols for SMART-Seq and Nextera XT, FACs, and tissue preparation and analysis & clustering links.
SOPs: RNA-Seq mouse whole cortex & hippocampus
SOPs: RNA-seq human multiple cortex areas
2. To monitor for a consistent, high-quality sampling of single-cell and single-nucleus RNA-Seq data, we have controls used in each application sample. These controls for mouse and human data are available to download for you to use in your own experiments.
The Cell Type Knowledge Explorer is a scientific and educational tool for exploration of human, marmoset, and mouse primary motor cortex cell types and the features that make them distinct.
Explore the Cell Type Knowledge Explorer
1. Learn from the scientists behind our Cell Types Knowledge Explorer about why it is made, key findings, and a walkthrough of the tool itself.
2. Read the science behind the cell type knowledge included in the Cell Type Knowledge Explorer in peer-reviewed publications.
Mouse Patch-seq (Scala et al 2021)
Mouse transcriptomics and epigenetics (Yao et al 2021)
Aligning cell types across species (Bakken et al 2021)
3. These use cases were designed to show researchers how to use the mouse data in the Cell Type Knowledge Explorer for their own research questions. The case titled “Experimental Design” shows how the Cell Type Knowledge Explorer can be used to guide research questions.
4. Python code for you to your own generate data visualizations that was used in the Cell Type Knowledge Explorer.
Replicate our data visualization
The Allen Brain Cell (ABC) Atlas provides a platform for visualizing multimodal single cell data across the mammalian brain and aims to empower researchers to explore and analyze multiple whole-brain datasets simultaneously.
1. From our product managers of ABC Atlas, here is a user guide to help you navigate all the features of ABC Atlas.
2. The ABC Atlas is under active development! See all the updates & patches to ABC Atlas on the Allen Brain Map Community Forum post that is updated regularly.
ABC Atlas channel on the Community Forum
3. Read the science behind some of the data sets included in the ABC Atlas in peer-reviewed publications.
Mouse Whole Brain (Yao et al 2023)
Mouse Whole Brain (Zhang et al 2023)
Human Whole Brain (Siletti et al 2023)
Human Alzheimer's disease (Gabitto et al 2021)
4. These use cases were designed to show researchers how to use the Whole Human Brain data in the ABC Atlas for their own research questions. The two use cases titled “Experimental Design” show how the ABC Atlas can be used to guide research questions, while the use case titled “Scientific Knowledge” shows how the ABC Atlas can be used studying and/or writing a literature review. The “coding” in the third use case shows users how to access the raw data for the Whole Human Brain using Jupyter notebooks.
Use Case: Scientific Knowledge
Use Case: Experimental Design with Coding
5. Learn from the scientists behind the new collection of studies from the BRAIN Initiative Cell Atlas Network (biccn.org) and published in Nature on Dec 14, 2023.
Whole Mouse Brain Paper Package Highlights Webinar
Related Collection of Scientific Publications
6. List of the 500 Gene panel used in the whole brain mouse. This is from the supplemental table 6 in Yao, et al. 2023. Also, the set of genes where expression in the spatial transcriptomics data is estimated (or "imputed") using information from the single cell RNA-seq data.
Whole Mouse Brain, Spatial Gene Panel List
Whole Mouse Brain, Spatial Imputed Genes
7. Tables including detailed information for each cluster, including neurotransmitter information, overall marker genes, main dissection region, and more. The first two buttons correspond to clusters in whole mouse brain (supplemental table 7 in Yao et al 2023) as published, and updated to include a comparison with clusters in a previous study of mouse cortex and hippocampus (Yao et al 2023). The third button relates to subclusters in whole human brain (unpublished supplemental information from Siletti et al 2023).
Whole Mouse Brain, Published Cluster Annotation Table
Whole Mouse Brain, Extended Cluster Annotation Table
Whole Human Brain, Subcluster Annotation Table
8. Download Excel files of the acronyms found in ABC Atlas with their corresponding full name, type of acronym (“types”), and the identifiers (“primary identifier”, “secondary identifier”, and “tertiary identifier”). See the table below for list of the types and identifiers used in the Excel file. Note that acronyms are now defined directly in the ABC Atlas in 'nomenclature cards'.
Whole Mouse Brain Cluster Annotations
Whole Mouse Brain Anatomical Annotations
Whole Human Brain Cluster Annotations
An ever-growing number of data sets are included in the ABC Atlas. This table lists the names, relevant publication, names and number of data points for every data sets available on the ABC Atlas (as of September 2025).
MapMyCells allows you discover what cell types your transcriptomics and spatial data corresponds with by comparing your data to our massive, high-quality reference datasets.
1. Learn from one of our scientists behind MapMyCells on how to use it, what type of data it accepts, what taxonomies are built of fit, which algorithms to select, and peek under the hood on how it works.
2. MapMyCells needs cell by gene matrix where rows are “cells” and columns are “genes”, which need to be in either a csv, csv.gz, or h5ad file format. This guide provides more details about how to prepare your file, including input file limits.
Input creation, file requirements, and limits
3. Unsure what algorithm or taxonomy to select in MapMyCells? These guides are for you.
4. The online version of MapMyCells will provide accurate cell type assignments for user-inputted data in most cases. However, there are some situations when using the direct scripts may be more appropriate: (1) if the reference taxonomy you are interested in is not one currently included in MapMyCells, (2) if the data set you have is quite large, (3) if you'd like to include these algorithms as part of an analysis pipeline, or (4) if you need to select a different set of genes for mapping (e.g., for mapping MERFISH data).
Run MapMyCells in python (cell_type_mapper)
Run MapMyCells in R (scrattch.mapping)
Most of these resources are part of the Brain Initiative Cell Census Network (BICCN) and/or Brain Initiative Cell Atlas Network (BICAN). The Allen Institute serves as the coordinating member of this network.
Discover all pages in this series:
This first section "What is..." lists introductory information on cell type topics.
Here are Allen Institute resources to help you understand the fundamentals of cell types and the research methods used to study cell types; including Patch-Seq, taxonomies, UMAPs, and more. These sections are designed for those who are unfamilar with cell type topics to help introduce them to the concepts.
Cells within a type exhibit similar structure and function that are distinct from cells in other types.
1. Learn from our Executive Vice President & Director of Brain Science, Hongkui Zeng, on the overview of cell types from the roots in evolution & development to the approaches in how to characterize, while also providing a roadmap for the future.
2. Cell Types 101 Webinar focuses on an introduction to the study of cell types! It covers evolving cell type definitions, cell types across species, the types of data used to define cell types, and the importance of having standard definitions for cell types.
The study of RNA expression (the transcriptome) and how it differs between cells / tissues / conditions.
This webinar is tutorial on Allen Cell Types Database. Starting at 4:13, we give an overview of single cell or nucleus transcriptomics. (Note: Webinar is from 2021; the taxonomies presented in the webinar are not the latest taxonomies from the Institute, as newer taxonomies have been created since 2021.)
This guide book describes how 10x single cell sequencing works with easy-to-understand language and graphics.
A cell type taxonomy is a specific analysis organizing cells into groups (or types), applied to a specific set of data, and saved in a standard format. Cell types are annotated with data-driven and historical information about their characteristics.
1. This user-friendly tutorial goes step by step on how cell type taxonomies are created, with linking to data from Allen Cell Types Database in the end. This is a great web-interface that is perfect for students to experts to begin exploring what a taxonomy is.
Explore the Interactive Walkthrough
Related tool: Allen Cell Types Database
Related Publication (Tasic et al 2016)
2. 'What is a taxonomy?' webinar is focused on the systematic classification of cell types and their hierarchical relationships. Much like species taxonomy (family, genus, species, etc.), researchers at the Allen Institute and their collaborators are working to create a standard taxonomy for cell types.
Uniform Manifold Approximation and Projections (or 'UMAPs') are helpful ways of displaying many types of data and are often referred to as one type of dimensionality reduction tool.
A UMAP is a common way to visualize cell types taxonomies. Learn how to interpret and analyze these graphs in this user guide.
Patch-Seq is modified version of Patch-Clamp, with the additional steps of extracting the nucleus to obtain transcriptomics and preserving the cell body to obtain morphology.
1. Patch-Seq was developed around late 2010s, with Allen Institute helping to optimize the technique. Here is a 2016 Allen Institute team talk giving an overview of Patch-Seq.
Watch the Team Talk
See other Patch-seq resources and publications
2. This webinar from 2024 gives an overview of what Patch-Seq is and how that data is used in the Cell Type Knowledge Explorer tool.
Watch the webinar
Explore motor cortex cell types in the Cell Type Knowledge Explorer
Discover all pages in this series: