Target Identification & Compound Profiling
Once the regulatory and conceptual landscape of pharmacological research has been established, the practical research pathway begins with a deceptively simple question: which biological molecule should the drug act upon? Phase 2 addresses target identification, the classification of druggable target families, the technological armamentarium used to identify and validate a target, and the chemical profiling techniques — structure-activity relationship analysis and ADMET prediction — used to characterise candidate compounds before they proceed to biological testing.
A drug target is the specific biological macromolecule — most commonly a protein, but occasionally a nucleic acid or lipid — with which a drug molecule interacts to bring about its characteristic pharmacological effect. The concept of a discrete molecular target, now central to rational pharmacology, replaced the earlier, largely empirical view of drug action as an unexplained whole-organism response, and it is this conceptual shift that allows modern drug discovery to be approached as a rational engineering problem rather than a matter of chance observation. A well-chosen target satisfies several criteria: it must be causally linked to the disease process (rather than merely correlated with it), it must be experimentally 'druggable' — that is, physically and structurally amenable to modulation by a small molecule, peptide, or biologic — and it must be sufficiently selectively expressed or engaged that modulating it does not produce unacceptable effects on unrelated physiological processes.
Target identification is the process of discovering the biological macromolecule — a protein, enzyme, receptor, ion channel, or nucleic acid — whose modulation is expected to produce the desired therapeutic effect. It represents the first step in rational, hypothesis-driven drug discovery, as distinct from purely phenotypic or serendipitous discovery. Contemporary target identification integrates genomics, proteomics, bioinformatics, and phenotypic screening to both discover candidate targets and validate their disease relevance before medicinal chemistry effort is committed to them.
Target validation is the process of establishing that modulating a candidate target actually alters the disease phenotype in a therapeutically favourable direction, and is as important scientifically as identification itself. Common validation approaches include genetic knockdown (using RNA interference or antisense oligonucleotides to transiently reduce target expression), genetic knockout (using CRISPR-Cas9 or older recombination-based techniques to eliminate the gene entirely, typically in an animal model), and pharmacological probe studies, in which a well-characterised tool compound is used to modulate the target acutely and observe the resulting phenotype. A poorly validated target is among the most common causes of late-stage clinical failure: a drug can be highly potent and selective against its intended target and still show no clinical benefit if that target was never truly disease-driving in humans, a phenomenon sometimes termed the 'right drug, wrong target' failure mode. For this reason, contemporary drug-discovery organisations increasingly demand convergent evidence from human genetics, disease-tissue expression data, and pharmacological probe studies before committing substantial resources to a target.
Druggable targets fall into a small number of structural and functional families, and recognising which family a candidate target belongs to immediately suggests the appropriate assay technology, the likely mode of pharmacological action, and relevant precedent drugs.
GPCRs constitute the largest single class of drug targets in current pharmacopoeia. These seven-transmembrane-domain proteins transduce extracellular signals — hormones, neurotransmitters, sensory stimuli — into intracellular second-messenger cascades via heterotrimeric G-proteins. Drugs acting at GPCRs may be agonists, antagonists, or allosteric modulators; classic examples include beta-blockers acting at adrenoceptors, opioids acting at the mu-opioid receptor, and antihistamines acting at the H1 receptor.
Ion channels may be ligand-gated (for example GABA-A or nicotinic acetylcholine receptors) or voltage-gated (sodium, potassium, or calcium channels). Drugs modulating ion channels underlie anaesthetics, antiepileptics, and antiarrhythmics, acting by altering channel open probability, conductance, or gating kinetics.
Enzymes are catalytic proteins that can be inhibited to block a specific biochemical pathway. Cyclooxygenase (COX) inhibition underlies NSAID action, angiotensin-converting enzyme (ACE) inhibition underlies antihypertensive therapy with captopril-class drugs, HMG-CoA reductase inhibition underlies statin therapy, and kinase inhibition (for example imatinib) underlies much of modern targeted oncology. Carbonic anhydrase is a further classical enzymatic target.
Nuclear receptors are ligand-activated transcription factors that, upon binding a lipophilic ligand, translocate to the nucleus and regulate gene expression directly. Glucocorticoid, thyroid hormone, retinoic acid, and PPAR receptors (targeted by glitazone antidiabetic drugs) belong to this family, and their pharmacology is characterised by relatively slow onset but long duration of action, reflecting the time required for transcriptional and translational responses.
Solute carrier (SLC) and ATP-binding cassette (ABC) transporters move ions and small molecules across membranes. Selective serotonin reuptake inhibitors (SSRIs) act by blocking the serotonin transporter (SERT); SGLT2 inhibitors such as canagliflozin block renal glucose reabsorption; and P-glycoprotein (MDR1), an ABC transporter, is central to multidrug resistance in cancer chemotherapy.
Kinases catalyse phosphorylation-based signal transduction and represent a major class of contemporary anticancer targets. Tyrosine kinase inhibitors such as imatinib and gefitinib, and serine/threonine kinase inhibitors, act by competing at the ATP-binding site or by allosteric inhibition of catalytic activity.
Nucleic acids themselves may be direct drug targets, particularly in oncology. Intercalating agents such as doxorubicin insert between DNA base pairs, alkylating agents such as cyclophosphamide form covalent DNA adducts, and topoisomerase inhibitors block DNA unwinding during replication.
Cytoskeletal and cell-wall structural proteins are also druggable. Tubulin is targeted by paclitaxel and the vinca alkaloids, disrupting microtubule dynamics during mitosis, while bacterial cell-wall synthesis is targeted by beta-lactam antibiotics and vancomycin.
Modern target identification typically follows a structured, multi-step workflow that moves from population-level genetic evidence to molecular-structural confirmation.
Genome-Wide Association Studies compare the genomes of affected and unaffected individuals to identify genetic variants statistically associated with disease, nominating the genes harbouring or nearest to those variants as candidate targets.
Transcriptomic profiling by RNA sequencing compares gene expression between diseased and healthy tissue, with consistently upregulated genes flagged as candidate targets. Proteomic techniques such as two-dimensional polyacrylamide gel electrophoresis (2D-PAGE) and mass spectrometry identify differentially expressed proteins and map protein–protein interaction networks that may reveal druggable nodes.
In parallel or as an alternative to target-based approaches, compounds may be screened directly for a desired cellular or organismal phenotype without prior knowledge of the molecular target, with target identity deconvoluted afterwards.
A compound of interest, immobilised on a solid support, is used to 'fish out' the proteins that bind it from a cell lysate, directly identifying its molecular binding partner(s).
X-ray crystallography and cryo-electron microscopy (Cryo-EM) resolve the three-dimensional structure of the candidate target, often in complex with a ligand, enabling structure-based drug design.
Candidate targets and their interactions are cross-referenced against curated databases: UniProt for protein function and sequence annotation, STRING for protein–protein interaction networks, ChEMBL for bioactivity data, and DrugBank for known drug–target relationships.
Once a hit compound has been identified, medicinal chemists undertake structure-activity relationship (SAR) analysis: the systematic modification of chemical structure to understand how specific structural features govern biological activity. SAR analysis identifies the pharmacophore — the minimum constellation of structural features essential for activity, such as a hydrogen-bond donor or acceptor, a hydrophobic region, or a charged group — and guides bioisosteric replacement, in which a functional group is substituted with one of similar electronic or steric character to improve selectivity, potency, or ADMET properties without abolishing activity.
Lipinski's Rule of Five
A widely used heuristic for predicting oral bioavailability, Lipinski's Rule of Five states that a drug-like molecule generally has a molecular weight below 500 Da, a calculated logP (lipophilicity) below 5, fewer than 5 hydrogen-bond donors, and fewer than 10 hydrogen-bond acceptors. Compounds violating two or more of these criteria are statistically less likely to be orally bioavailable, although numerous clinically important exceptions exist (particularly among natural products and biologics), so the rule should be applied as a guide rather than an absolute filter.
ADMET — Absorption, Distribution, Metabolism, Excretion, and Toxicity — prediction is undertaken early in compound profiling to eliminate candidates likely to fail for pharmacokinetic or safety reasons, long before the expense of animal testing is incurred. Each ADMET parameter has established in-silico and in-vitro surrogate measures.
Oral absorption is estimated from calculated logP (an ideal range of approximately 1–3), pKa, aqueous solubility, and Caco-2 cell monolayer permeability, with a permeability coefficient above roughly 1 × 10⁻⁶ cm/s generally considered indicative of good intestinal absorption. MDCK (Madin–Darby Canine Kidney) cell monolayers serve as a complementary permeability model.
Distribution is characterised by plasma protein binding (expressed as the unbound fraction, fu), distribution coefficient (logD), and polar surface area — a PSA below approximately 90 Ų is generally required for central nervous system penetration. Volume of distribution (Vd) values of 0.1–1 L/kg suggest predominantly plasma confinement, while values exceeding 5 L/kg indicate extensive tissue distribution.
Hepatic metabolism is assessed through cytochrome P450 (CYP450) substrate and inhibitor profiling across the major isoforms — CYP3A4, CYP2D6, CYP2C9, CYP2C19, and CYP1A2 — typically using human liver microsomes (HLM), from which intrinsic metabolic half-life is derived. CYP-mediated drug–drug interaction risk is a major cause of late-stage attrition and post-marketing withdrawal.
Renal clearance, biliary excretion, and P-glycoprotein-mediated efflux (commonly assessed via the Caco-2 basolateral-to-apical/apical-to-basolateral, or B-A/A-B, permeability ratio) together determine the elimination pathway and half-life of a compound.
Computational toxicity prediction platforms such as DEREK Nexus and Sarah Nexus flag structural alerts associated with known toxicophores. Of particular concern in early profiling is hERG channel liability, a predictor of cardiac QT-interval prolongation and potentially fatal arrhythmia; AMES mutagenicity prediction and blood-brain barrier (BBB) penetration scoring are likewise routinely generated at this stage.
Widely used ADMET prediction platforms include SwissADME (a free, web-based tool), pkCSM, ADMETlab 2.0, and commercial packages such as Schrödinger's QikProp, MOE, and Maestro Glide, the latter also supporting molecular docking against resolved target structures.
In-silico (computational) screening allows enormous virtual compound libraries — commonly tens of millions of commercially available or synthetically accessible structures — to be evaluated computationally against a target of interest before a single physical compound is purchased or synthesised. Virtual screening approaches are broadly divided into structure-based methods, which require a resolved or homology-modelled three-dimensional structure of the target, and ligand-based methods, which instead exploit the structural features of known active compounds (via pharmacophore modelling or quantitative structure-activity relationship, QSAR, analysis) to prioritise structurally similar candidates from a virtual library. The principal advantage of in-silico screening is economy: it allows a research programme to concentrate scarce synthetic and biological screening resources on the small subset of compounds most likely to show genuine activity, dramatically improving hit rates relative to purely random physical screening.
Molecular docking computationally predicts the preferred orientation (or 'pose') of a small molecule within the binding site of a target protein, and estimates the strength of that interaction using a scoring function that accounts for hydrogen bonding, hydrophobic contact, electrostatic complementarity, and steric fit. Docking proceeds in two conceptually distinct steps: a search algorithm generates a large number of candidate poses for each ligand within the defined binding site, and a scoring function then ranks these poses (and, across compounds, ranks different ligands against one another) by predicted binding affinity. Widely used docking platforms include AutoDock Vina (freely available), Glide (part of the Schrödinger Maestro suite), and GOLD. Because scoring functions remain imperfect approximations of true binding thermodynamics, docking results are conventionally treated as a hypothesis-generating filter rather than a definitive prediction, and top-ranked compounds are always confirmed experimentally using the receptor-binding and functional assays described in Phase 3.
Bioinformatics — the application of computational tools to biological data — underlies almost every stage of modern target identification and compound profiling. Sequence-analysis tools identify homologous proteins across species, informing the choice of an appropriate animal model for later in-vivo testing (Phase 4). Structural bioinformatics tools, including homology modelling algorithms such as MODELLER and, more recently, deep-learning-based structure prediction tools such as AlphaFold, generate three-dimensional target models even in the absence of an experimentally resolved crystal structure, considerably expanding the range of targets amenable to structure-based drug design. Pathway and network-analysis tools situate a candidate target within its broader signalling context, helping to anticipate on-target and off-target consequences of pharmacological modulation. Together, these bioinformatic resources allow a research team to make evidence-based prioritisation decisions long before committing to the expense of wet-laboratory screening.
Phase 2 has traced the path from an unvalidated biological hypothesis to a chemically and pharmacokinetically profiled lead compound: the classification of druggable target families, the genomic-to-structural workflow of target identification, and the SAR and ADMET techniques used to refine a hit into a viable candidate. With a validated target and a profiled compound in hand, the next stage of the pipeline — described in Phase 3 — is to test that compound's actual biological activity using in-vitro pharmacological screening.