More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Your genome and a large language model’s weights both boil down to long symbol strings that gain purpose only when “read” by another system. In DNA’s case, enzymes unzip and transcribe the code, and ribosomes translate it into proteins. In neural networks, the raw weights sit inert until the inference engine multiplies inputs through layers, producing predictions. Neither sequence “contains” a mind or instructions in the conventional sense—they’re inert until an external process interprets them.
Both sequences are digital, discrete codes designed for reliable copying. In evolution, natural selection ran an unbroken search over roughly four billion years and countless organisms, filtering out every variation that didn’t boost reproductive success. The surviving DNA is the compressed record—about three billion letters, or 750 MB—that captures the regularities rewarded by that process. Likewise, gradient descent tweaks billions of numerical weights a hair at a time, over trillions of examples drawn from almost everything humanity has written. What remains in the model is not the training texts themselves but the statistical residue of being “wrong” and corrected trillions of times.
Neither the genome nor the weights function as blueprints. They’re more like programs written in an alphabet—ATGC for life, token IDs for language models—that, when run in the right environment, produce complex results. You won’t find a single gene for your hand any more than you’ll find one weight storing “Paris is the capital of France.” Knowledge lives in the distributed process, not in isolated entries. Both evolution and gradient descent deliver genuine competence—eyes, code generation, even math problem solving—without any single locus of comprehension.
Questions about this article
No questions yet.