Please use this identifier to cite or link to this item: http://localhost:8081/jspui/handle/123456789/21313
Full metadata record
DC FieldValueLanguage
dc.contributor.authorChauhan, Deepanshu-
dc.date.accessioned2026-08-07T11:43:03Z-
dc.date.available2026-08-07T11:43:03Z-
dc.date.issued2023-06-
dc.identifier.urihttp://localhost:8081/jspui/handle/123456789/21313-
dc.guideGupta, Manuen_US
dc.description.abstractWe study deep contextual bandits, a class of contextual bandits where each context-arm pair is associated with a feature vector having an unknown reward-generating function. We compared already available algorithms like LinUCB, Neural UCB, Neural LinUCB, and Neural Linear. Implemented Neural LinUCB algorithm on the real-world dataset, which has finite action, and modified the algorithm for infinite action cases to optimize an unknown polynomial function. Our algorithm does not use Deep Neural Network to optimize the unknown reward function, as in the case of Neural LinUCB; instead, it relies on the reinforcement learning algorithm only. The objective of algorithm is to learn the maxima of the reward generating function which is a polynomial but degree of polynomial is unknown to the agent. The proposed algorithm is able to learn reward generating function in most of the cases.en_US
dc.language.isoenen_US
dc.publisherIIT Roorkeeen_US
dc.titleOnline Learning for Noisy Polynomialen_US
dc.typeDissertationsen_US
Appears in Collections:MASTERS' THESES (MFSDS & AI)

Files in This Item:
File Description SizeFormat 
21566005_Deepanshu Chauhan.pdf1.56 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.