Product Experimentation with Doubly Robust Estimation: When Both Your Models Are Wrong in LLM Applications