
This course provides a comprehensive introduction to Cheminformatics, an interdisciplinary field that combines chemistry, computer science, and data science to represent, organize, analyze, and interpret chemical information using computational tools. Designed in accordance with the UGC SWAYAM framework, the curriculum is organized into two complementary units that guide students from the fundamentals of digital molecular representation to modern computer-aided drug discovery techniques.
The first unit establishes the foundation of cheminformatics by introducing its scope, importance, and role in modern chemical research. Students learn how molecules are represented digitally using widely adopted formats such as SMILES, InChI, PDB, SDF, and MOL2, and gain practical experience in representing, reading, writing, and manipulating molecular structures using the RDKit toolkit. The unit also introduces essential molecular editing operations, including hydrogen addition/removal, structural modification, conformer generation, and substructure matching. Finally, students explore major public chemical databases such as PubChem, ChemSpider, and ZINC, developing the skills required to efficiently search, retrieve, and manage chemical information from large molecular repositories.
The second unit focuses on the computational methods widely used in computer-aided drug discovery (CADD). Students learn how to calculate molecular descriptors, fingerprints, and physicochemical properties that quantitatively describe molecular structures and influence their biological behavior. The course further explores molecular similarity analysis, substructure searching, and virtual screening techniques for identifying promising lead compounds from large chemical libraries. It also introduces ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity)prediction using tools such as SwissADMET, enabling students to evaluate drug-likeness and bioavailability in the early stages of drug design. The unit concludes with an introduction to protein structure databases and molecular docking using SwissDock, followed by hands-on projects involving substructure matching, similarity analysis, and interpretation of computational results.
By the end of this course, students will be able to represent molecules digitally, retrieve and analyze chemical information from public databases, compute molecular descriptors and fingerprints, perform molecular similarity searches and virtual screening, predict drug-like properties using ADMET tools, and carry out introductory molecular docking studies. These foundational computational skills prepare learners for advanced studies in cheminformatics, computational chemistry, medicinal chemistry, and modern data-driven drug discovery.
- Teacher: Prashant Kumar Gupta