Production
How it works
A construct — the gene of interest, typically with an affinity tag and any fusion partner needed for solubility or detection — is introduced into the expression system best matched to the target protein. Bacterial (E. coli) expression is fast, cheap and gives high yields of simple, non-glycosylated proteins, enzymes and fragments, but its cytoplasm lacks the machinery for complex folding and post-translational modification, and high-level expression there often drives the protein into insoluble aggregates called inclusion bodies rather than its correctly folded, soluble form. Mammalian expression is slower and more expensive, but produces properly folded, human-like glycosylated protein — the default choice for antibodies and other complex glycoproteins where post-translational modification affects function or immunogenicity.
What you measure
Deliverables from a production run are assessed on yield (typically mg of protein per litre of culture), solubility (soluble supernatant versus insoluble inclusion-body fraction), and identity/purity at the crude stage — usually by SDS-PAGE and, where relevant, a quick activity or binding check — before the material moves on to purification. For a construct expressed for the first time, a small-scale expression trial across a few conditions is standard practice before committing to a larger preparative-scale run.
A typical experiment
For E. coli, expression is typically induced with IPTG once the culture reaches mid-log growth, and post-induction temperature, induction time and inducer concentration are the main levers for balancing yield against solubility — lowering the temperature to roughly 16–25 °C after induction is one of the simplest ways to reduce inclusion-body formation, at some cost to overall yield, and fusion tags such as MBP or GST can further improve solubility for a difficult construct. Mammalian expression instead runs on a timescale of days to weeks in suspension culture, with the secreted protein harvested from the medium — convenient for downstream purification, since the bulk of host-cell protein stays inside the cells.
Applications
Custom constructs for in-house assays. Production feeds material directly into the platform's own ITC, AUC, SPR, DSC, nanoDSF and DLS assays, so the construct, tag position and buffer can be planned together with the downstream measurement in mind.
Difficult or novel targets. For proteins not available from any commercial catalogue — new mutants, truncations, fusion constructs or labelled variants — in-house expression is often the only route to material.
Antibody & complex glycoprotein production. Mammalian expression supplies correctly folded, glycosylated protein for antibody fragments and other targets where post-translational modification affects activity or stability.
Small-scale screening before scale-up. Testing expression conditions at small scale before committing to a larger preparative run reduces the risk of discovering a solubility problem only after a large culture has already been grown.
Strengths & limitations
The central trade-off in bacterial expression is yield versus solubility: conditions that favour high yield (higher temperature, stronger induction) tend to push the target protein into inclusion bodies, while conditions that favour soluble, correctly folded protein tend to reduce overall yield — so the optimal condition is usually found empirically for each construct rather than assumed. If a target only expresses as inclusion bodies, it can sometimes still be recovered by solubilization and refolding, but that adds a step and is not guaranteed to succeed, particularly for larger or disulfide-rich proteins. Choosing mammalian expression sidesteps most of this at the cost of longer timelines and lower yields per litre, so the choice of system is really a choice about which constraint — time, yield, or folding fidelity — matters most for a given project.
Frequently asked questions
What do I need to give you to get started?
The amino acid sequence of what you want, in a form we can read — FASTA, a UniProt accession with the boundaries you want, or a plasmid map. If you already have an expression construct, send the plasmid with its sequence and the strain or cell line it was made for.
Beyond the sequence: how much protein you need and at what purity, what it is for (crystallography, an assay, an immunisation, a binding measurement — each implies different requirements), whether a tag can stay on, and what buffer it should end up in.
On the administrative side we need the biosafety classification of the material, and a material transfer agreement if the construct comes from a third party. Toxin genes, sequences from pathogens and anything requiring containment above level 1 have to be declared at the start.
Which expression system should I use, and how long does each take?
E. coli is the first thing to try for anything bacterial, for isolated domains, for proteins without disulfides in non-native positions and for anything that needs isotope labelling. From an existing plasmid to purified protein is one to three weeks.
Insect cells with baculovirus handle larger eukaryotic proteins and multi-subunit complexes, and tolerate constructs that E. coli refuses outright. The timeline is longer because the virus has to be generated and amplified: four to eight weeks from gene to protein, of which most is virus work.
Mammalian transient expression in HEK293 is the route for secreted proteins, glycoproteins where human-type glycosylation matters, and antibodies. Three to five weeks from a gene, less from an existing plasmid. Stable cell lines take months and only pay off for repeated large-scale production.
The honest version: for a protein nobody has expressed before, the system is chosen by trying two in parallel at small scale rather than by reasoning from the sequence. That costs two weeks and saves considerably more.
What yield can I expect?
Anywhere from nothing to hundreds of milligrams per litre, and nobody can predict which before trying.
Order-of-magnitude figures: a well-behaved soluble protein in E. coli gives 5-50 mg per litre of culture and occasionally far more; a secreted protein transiently expressed in HEK293 typically gives 5-50 mg/l; baculovirus-expressed intracellular proteins commonly land at 1-10 mg/l. Membrane proteins in any system are an order of magnitude below that.
The variation between targets is larger than the variation between systems, and two constructs of the same protein differing only in their boundaries can differ tenfold in yield. This is why a quotation is written against a defined amount of work — a number of litres, a number of constructs, a number of purification steps — rather than against a guaranteed quantity of protein. If you need a firm amount, the sensible route is a small-scale test first and a scale-up decision afterwards.
Which tag should I use, and will you remove it?
His6 is the default: small, works in every system, and IMAC is a robust first step. Its weaknesses are that it binds other metal-binding proteins from the host and that it is useless if the lysate contains chelators.
Strep-II gives cleaner one-step material at lower capacity. GST and MBP are larger and often improve solubility, which is their real purpose; MBP in particular rescues proteins that are otherwise entirely in inclusion bodies. SUMO fusions cleave to leave a native N-terminus, which matters when the first residue is part of the function.
Removal is done with a site-specific protease — TEV, 3C/PreScission, thrombin — followed by a subtractive IMAC step to take out the cut tag and the protease. Expect to lose 20-40% of the material across cleavage and re-purification, and expect a small scar of one or a few residues depending on the protease. TEV leaves a glycine-serine, 3C leaves a GP.
Whether to remove it depends on the downstream use. For an SPR surface a tag is an asset; for crystallography it is usually removed; for a functional assay it depends on whether the terminus is involved.
Do you test at small scale before scaling up?
For any target that has not been produced before, yes, and it is the step that decides whether the project works.
Small-scale screening means several millilitres of culture per condition, tested in parallel: two or three strains, two induction temperatures, two media, and where relevant several construct boundaries. Solubility is read from a lysate on a gel or from a small-scale IMAC, and the whole thing takes a few days.
Testing construct boundaries is what changes outcomes most often. A domain that starts three residues into a helix expresses as an insoluble mess; the same domain with boundaries taken from a structure or a disorder prediction behaves. Screening four to eight boundary variants at small scale is cheap next to scaling up the wrong one.
The same logic applies to the purification buffer. A nanoDSF or DLS screen on the small-scale material identifies the conditions where the protein is stable, before a litre of culture is committed.
What quality control comes with the protein?
As standard: SDS-PAGE of the final material with an estimate of purity, concentration by A280 with the calculated extinction coefficient, and the chromatograms from every purification step.
On request, and worth having for most uses: intact mass by mass spectrometry, which confirms identity, tag cleavage and the absence of unexpected modification; analytical SEC or DLS for the oligomeric state and aggregate content; and peptide mapping if the identity needs to be established beyond the mass.
Endotoxin measurement and removal apply to material destined for cell assays or animal work, and have to be requested from the start because they change how the purification is run.
What is not included unless it is written into the order is a functional or activity assay. Purity is not activity — a protein can be 98% pure by gel and entirely dead — and only you have the assay that tells the difference. If you have one, running it on a small aliquot before scale-up is time well spent.
In what form will I receive the protein?
In aliquots, flash-frozen in liquid nitrogen and stored at -80 °C, in the buffer agreed in the order. Aliquot size is set by how you will use it, because a tube that gets thawed five times is a different sample by the fifth time.
Concentration is whatever the protein tolerates. Some proteins go to 20 mg/ml without complaint, others aggregate above 1 mg/ml, and pushing past that limit to hit a requested number does more harm than delivering a larger volume of a more dilute solution. If your application needs a specific concentration, say so early so the final step can be designed around it.
Shipping is on dry ice for frozen material. Local collection is straightforward. For a long international shipment, lyophilisation is sometimes the better option, though it should be tested on an aliquot first.
You also get the documentation: sequence of what was actually made, purification method, chromatograms, gels, concentration, buffer composition, and the storage recommendation.
Can you produce membrane proteins, glycoproteins or difficult targets?
Glycoproteins and secreted proteins are handled in HEK293 or insect cells, with the caveat that glycosylation is heterogeneous and differs between systems. If a defined glycoform matters, that constrains the cell line and sometimes requires glycosylation-deficient variants or enzymatic trimming afterwards.
Membrane proteins are possible and are a different scale of project. Expression levels are low, a detergent screen is needed to find one that extracts and stabilises the protein, and the purification usually loses most of the starting material. Budget for months rather than weeks, and for a real chance of failure. A small-scale expression and detergent screen before any commitment is not optional here.
Intrinsically disordered proteins express well and purify strangely — they run high on SEC, look wrong on gels and are protease-sensitive, so heat treatment and protease inhibitors are part of the standard approach.
Toxic proteins, proteins requiring a specific cofactor and multi-subunit complexes needing co-expression are all doable, and all need discussion before a quotation makes any sense.
What happens if the protein does not express, or comes out insoluble?
This is common enough that it should be planned for rather than treated as an accident. Roughly half of eukaryotic targets fail on the first attempt in E. coli.
The usual sequence of responses: lower the induction temperature to 16-20 °C overnight, change strain to one handling rare codons or disulfides, change the tag to MBP or SUMO, redraw the construct boundaries, co-express chaperones or a binding partner, or move to another expression system. Each of these is a defined piece of work with its own cost, and they are tried in an order agreed with you rather than indefinitely.
Refolding from inclusion bodies is a last resort. It works for small proteins with simple topology and few disulfides, and it fails silently for everything else — you can recover soluble material that is not correctly folded, which is worse than recovering nothing.
The practical arrangement is a decision point after the small-scale screen: continue, change strategy, or stop. Structuring an order that way means you find out in three weeks rather than three months that a target needs a different approach.
Who owns the constructs and the results, and can I take part in the work?
The material and the data belong to you. Constructs made for your project are yours, and plasmid stocks are returned or kept for you as you prefer. Confidentiality covers sequences and results, and nothing is published or reused without your agreement.
Work done as a service does not by itself warrant authorship; a contribution that goes beyond that — designing a strategy, solving a target that needed real method development — is worth discussing when the paper is written, and the usual practice is to settle it at the start rather than at the end.
Taking part is welcome. Coming for the purification of your own protein is common, particularly for a target you will produce again yourself afterwards, and the small-scale screening stage is a good moment to be present because that is where the decisions are made. Full training on the production pipeline is a longer commitment and makes sense for people setting up the same work in their own laboratory.
Instruments
Assist Plus
Jasco UV-750